About StackOne: StackOne is the AI Integration Gateway for SaaS products and AI Agents. Backed by GV and Workday Ventures ($24M raised), we help builders of SaaS platforms and AI Agents orchestrate hundreds of scalable, accurate, and enterprise-grade integrations. Our platform combines 25,000 pre-mapped actions on 200 connectors, an AI-powered integration development toolkit, plus security by design: a real-time architecture, managed authentication and permissions, and end-to-end observability. Join us on our fast trajectory to build the future of agentic integrations. About the role We're looking for a Senior Platform Engineer to own how StackOne is built, shipped, and run, as we scale across our own cloud and into our customers' clouds. You'll own the infrastructure behind the platform, our deployment pipeline and developer tooling, and how we package StackOne to run inside customers' own AWS, GCP, or Azure accounts. It's a hands on role with broad scope. You write code and tooling, you own the IaC other engineers depend on, and you set the standard for how every new repository gets deployed and secured. You'll report directly to the CTO and work closely with our Security Engineer and tech leads. Responsibilities Own our infrastructure at scale: the AWS estate today (ECS Fargate, Aurora, ElastiCache, MSK, OpenSearch, Lambda, KMS) and the AWS CDK to Terraform migration. Keep it reliable, observable, and cost aware. Build out the deployment pipeline as we consolidate toward a monorepo: the CI/CD that ships every service, plus an automated end to end testing harness with incremental (affected only) testing that stays fast as we grow. Ship into customers' own clouds (self hosted / BYOC): the Terraform modules, container images, runbooks, and documentation for internal teams and customers. Own the release and upgrade path for self hosted customers: versioned, signed releases, a supported version policy, and a call home for usage and version reporting. Set the standard for new repositories: partner with the Security Engineer and tech leads so new projects (product, internal tools, and vibe coded prototypes) ship secure and deployable from day one. Templates and golden paths, not gate keeping. Raise reliability: SLOs, observability, and incident response. Make the system easy to operate when something breaks at 3am. Treat infrastructure as a product: paved roads and self serve tooling so product engineers ship without waiting on you. Use AI in the workflow: lean on LLMs and agents for IaC generation, test scaffolding, and runbook drafting, with guardrails you trust. What we're looking for 4+ years in platform, infrastructure, SRE, or DevOps, with hands on AWS at scale: ECS/Fargate, RDS/Aurora, Lambda, networking, IAM, KMS. Deep IaC ability: Terraform and/or AWS CDK, comfortable owning modules other teams build on. Strong CI/CD and monorepo build experience: caching and incremental/affected only test execution. You've made a slow pipeline fast. Strong coding in TypeScript, Python, or Go. You build tooling, not just YAML and configs. Containers in production (Docker); Kubernetes a plus. Security minded: you bake secure defaults into pipelines and repo templates, and partner with security rather than routing around it. A clear writer: docs and runbooks that internal engineers and external customers can follow. End to end ownership: you scope, ship, and measure, and automate instead of running manual checklists. Nice to have Shipping software into customers' own cloud or on prem (BYOC / self hosted), or a platform like Nuon, Replicated, or Omnistrate. Multi cloud (GCP or Azure) IaC, and exposure to Temporal, Kafka/MSK, ClickHouse, or OpenSearch. Our stack Cloud & infra: AWS (ECS Fargate, Aurora Postgres, ElastiCache, MSK, OpenSearch, Lambda, S3, KMS, CloudFront, WAF), Cloudflare (Workers, WAF) IaC: AWS CDK today, migrating to Terraform Data & messaging: Postgres, Redis, Kafka, OpenSearch, ClickHouse CI/CD: GitHub Actions Observability & analytics: Datadog, Sentry, Metabase Languages: TypeScript (Node.js), Python Benefits Meaningful share options (EMI) 25 days holiday + 1 additional day per year of tenure Private health insurance, including dental & optical £15/day London office lunch budget, up to £120/month £1,000 home office setup + £500/year top up Annual team offsite to sunny spots Join one of Europe's fastest growing startups Work with a veteran team of ex Google, Microsoft, Oracle, Coinbase, JP Morgan and more Health, fitness and gift card discounts; Cycle2Work and Electric Cars scheme London (hybrid, 2 days/week) preferred; open to remote within the UK We believe diversity drives innovation. We encourage individuals from all backgrounds to apply. As an equal opportunity employer, we celebrate diversity and are committed to creating an inclusive environment for all employees.
28/07/2026
Full time
About StackOne: StackOne is the AI Integration Gateway for SaaS products and AI Agents. Backed by GV and Workday Ventures ($24M raised), we help builders of SaaS platforms and AI Agents orchestrate hundreds of scalable, accurate, and enterprise-grade integrations. Our platform combines 25,000 pre-mapped actions on 200 connectors, an AI-powered integration development toolkit, plus security by design: a real-time architecture, managed authentication and permissions, and end-to-end observability. Join us on our fast trajectory to build the future of agentic integrations. About the role We're looking for a Senior Platform Engineer to own how StackOne is built, shipped, and run, as we scale across our own cloud and into our customers' clouds. You'll own the infrastructure behind the platform, our deployment pipeline and developer tooling, and how we package StackOne to run inside customers' own AWS, GCP, or Azure accounts. It's a hands on role with broad scope. You write code and tooling, you own the IaC other engineers depend on, and you set the standard for how every new repository gets deployed and secured. You'll report directly to the CTO and work closely with our Security Engineer and tech leads. Responsibilities Own our infrastructure at scale: the AWS estate today (ECS Fargate, Aurora, ElastiCache, MSK, OpenSearch, Lambda, KMS) and the AWS CDK to Terraform migration. Keep it reliable, observable, and cost aware. Build out the deployment pipeline as we consolidate toward a monorepo: the CI/CD that ships every service, plus an automated end to end testing harness with incremental (affected only) testing that stays fast as we grow. Ship into customers' own clouds (self hosted / BYOC): the Terraform modules, container images, runbooks, and documentation for internal teams and customers. Own the release and upgrade path for self hosted customers: versioned, signed releases, a supported version policy, and a call home for usage and version reporting. Set the standard for new repositories: partner with the Security Engineer and tech leads so new projects (product, internal tools, and vibe coded prototypes) ship secure and deployable from day one. Templates and golden paths, not gate keeping. Raise reliability: SLOs, observability, and incident response. Make the system easy to operate when something breaks at 3am. Treat infrastructure as a product: paved roads and self serve tooling so product engineers ship without waiting on you. Use AI in the workflow: lean on LLMs and agents for IaC generation, test scaffolding, and runbook drafting, with guardrails you trust. What we're looking for 4+ years in platform, infrastructure, SRE, or DevOps, with hands on AWS at scale: ECS/Fargate, RDS/Aurora, Lambda, networking, IAM, KMS. Deep IaC ability: Terraform and/or AWS CDK, comfortable owning modules other teams build on. Strong CI/CD and monorepo build experience: caching and incremental/affected only test execution. You've made a slow pipeline fast. Strong coding in TypeScript, Python, or Go. You build tooling, not just YAML and configs. Containers in production (Docker); Kubernetes a plus. Security minded: you bake secure defaults into pipelines and repo templates, and partner with security rather than routing around it. A clear writer: docs and runbooks that internal engineers and external customers can follow. End to end ownership: you scope, ship, and measure, and automate instead of running manual checklists. Nice to have Shipping software into customers' own cloud or on prem (BYOC / self hosted), or a platform like Nuon, Replicated, or Omnistrate. Multi cloud (GCP or Azure) IaC, and exposure to Temporal, Kafka/MSK, ClickHouse, or OpenSearch. Our stack Cloud & infra: AWS (ECS Fargate, Aurora Postgres, ElastiCache, MSK, OpenSearch, Lambda, S3, KMS, CloudFront, WAF), Cloudflare (Workers, WAF) IaC: AWS CDK today, migrating to Terraform Data & messaging: Postgres, Redis, Kafka, OpenSearch, ClickHouse CI/CD: GitHub Actions Observability & analytics: Datadog, Sentry, Metabase Languages: TypeScript (Node.js), Python Benefits Meaningful share options (EMI) 25 days holiday + 1 additional day per year of tenure Private health insurance, including dental & optical £15/day London office lunch budget, up to £120/month £1,000 home office setup + £500/year top up Annual team offsite to sunny spots Join one of Europe's fastest growing startups Work with a veteran team of ex Google, Microsoft, Oracle, Coinbase, JP Morgan and more Health, fitness and gift card discounts; Cycle2Work and Electric Cars scheme London (hybrid, 2 days/week) preferred; open to remote within the UK We believe diversity drives innovation. We encourage individuals from all backgrounds to apply. As an equal opportunity employer, we celebrate diversity and are committed to creating an inclusive environment for all employees.
About Us: Solirius Reply, part of the Reply Group, is a technology consultancy and digital transformation partner that helps organisations solve complex challenges through strategy, design, engineering, and delivery. We work closely with our clients to deliver secure, accessible, user-focused services that evolve with their needs. By combining deep technical expertise with people-centred design, we create solutions that deliver meaningful, lasting impact. Our consultants partner directly with client teams, embedding into organisations to understand their goals, challenges, and users. This collaborative approach enables us to deliver tailored solutions that drive measurable outcomes across public and private sectors. Past and present clients include the Ministry of Justice, Department for Education, Ministry of Housing, Communities and Local Government, UEFA, International Olympic Committee, and Mercedes-Benz. Our services span the full digital delivery lifecycle, including architecture, engineering, delivery management, user-centred design, business analysis, data, DevOps, and AI. We operate as a collaborative and inclusive organisation that empowers our people to take ownership, innovate, and develop their expertise. As an equal opportunities employer, we are committed to encouraging equality, diversity, and social mobility, while creating opportunities for our teams to work on meaningful projects that deliver lasting impact. About You: You are a motivated and adaptable professional with a strong analytical mindset and a passion for using technology to solve real-world problems. You enjoy working in collaborative, agile teams and take pride in delivering high-quality solutions that make a tangible impact. With strong communication skills and a consultative approach, you're comfortable engaging with clients, understanding their needs, and translating them into effective outcomes. The Role: We are looking for an experienced Senior Data Engineer to join our team here at Solirius.You will be working as part of a team, developing and delivering exciting projects with afantastic team of technology experts. You will be responsible for designing, developing, and maintaining data pipelines andsystems that enable data analysis and machine learning. You will also collaborate with datascientists, analysts, and other stakeholders to ensure data quality and reliability. Key Responsibilities: Develop Data Engineering solutions for our clients/projects Design and build data models, schemas to support business requirements Develop and maintain data ingestion and processing systems using various tools andtechnologies, such as SQL, NoSQL, ETL, Luigi, Airflow, Argo, etc. Implement data storage solutions using different types of databases, such asrelational, non-relational, or cloud-based. Working collaboratively with the client and cross-functional teams to identify andaddress data-related issues and opportunities. Stay updated with the latest trends and developments in the data engineering field,such as modern data stack, big data technologies, cloud computing, etc Assist with defining the processes needed to achieve operational excellence Ensure data quality across all projects/clients Define and manage SLA's for data sets and processes throughout production Lead on the design, build and launch of new data models and pipelines Key Skills/Experience: Experience in operating as part of data engineering teams and independently. Line/Team management experience, leadership experience Experience of working with cloud infrastructure (Azure or AWS, GCP is beneficial) SQL and relational databases (e.g. MS SQL/Azure SQL, PostgreSQL) You have framework experience within either Flask, Tornado or Django, Docker Experience working with ETL pipelines is desirable e.g. Luigi, Airflow or Argo Experience with big data technologies, such as Apache Spark, Hadoop, Kafka, etc. Data acquisition and development of data sets and improving data quality Preparing data for predictive and prescriptive modelling Hands on coding experience, such as Python Reporting tools (e.g. Tableau, PowerBI, Qlik) GDPR and Government Service Standard (desirable) Passionate, motivated and enthusiastic about developing technology solutions. Experience working in an Agile development environment Data architecture experience. Competitive Salary Bonus Scheme Private Healthcare Insurance 25 Days Annual Leave + Bank Holidays Up to 10 days allocated for development training per year Enhanced Parental Leave Paid Fertility Leave (5 Days) Statutory & Contributory Pension EAP with Gym Membership Benefits Cycle to Work and Electric Vehicle schemes Flexible Working Annual Away Days/Company Socials Diversity and Inclusion As an equal opportunities employer, we are committed to creating a work environment that supports, celebrates, encourages and respects all individuals, where all processes are based on merit, competence and business needs. Encouraging high social mobility is really important to us. We foster an inclusive culture by welcoming different perspectives, enabling equitable opportunities and promoting open dialogue. This commitment is reflected in initiatives such as our gender diversity group and our focus on mental health and wellbeing. Whatever stage you are at, you will find an environment where you can thrive. Should you require further assistance or require any reasonable adjustments to be put in place to better support your application process, please do not hesitate to raise this with us. As a Disability Confident employer, we are committed to ensuring our recruitment process is accessible and inclusive, enabling all candidates to demonstrate their skills, experience and potential.
27/07/2026
Full time
About Us: Solirius Reply, part of the Reply Group, is a technology consultancy and digital transformation partner that helps organisations solve complex challenges through strategy, design, engineering, and delivery. We work closely with our clients to deliver secure, accessible, user-focused services that evolve with their needs. By combining deep technical expertise with people-centred design, we create solutions that deliver meaningful, lasting impact. Our consultants partner directly with client teams, embedding into organisations to understand their goals, challenges, and users. This collaborative approach enables us to deliver tailored solutions that drive measurable outcomes across public and private sectors. Past and present clients include the Ministry of Justice, Department for Education, Ministry of Housing, Communities and Local Government, UEFA, International Olympic Committee, and Mercedes-Benz. Our services span the full digital delivery lifecycle, including architecture, engineering, delivery management, user-centred design, business analysis, data, DevOps, and AI. We operate as a collaborative and inclusive organisation that empowers our people to take ownership, innovate, and develop their expertise. As an equal opportunities employer, we are committed to encouraging equality, diversity, and social mobility, while creating opportunities for our teams to work on meaningful projects that deliver lasting impact. About You: You are a motivated and adaptable professional with a strong analytical mindset and a passion for using technology to solve real-world problems. You enjoy working in collaborative, agile teams and take pride in delivering high-quality solutions that make a tangible impact. With strong communication skills and a consultative approach, you're comfortable engaging with clients, understanding their needs, and translating them into effective outcomes. The Role: We are looking for an experienced Senior Data Engineer to join our team here at Solirius.You will be working as part of a team, developing and delivering exciting projects with afantastic team of technology experts. You will be responsible for designing, developing, and maintaining data pipelines andsystems that enable data analysis and machine learning. You will also collaborate with datascientists, analysts, and other stakeholders to ensure data quality and reliability. Key Responsibilities: Develop Data Engineering solutions for our clients/projects Design and build data models, schemas to support business requirements Develop and maintain data ingestion and processing systems using various tools andtechnologies, such as SQL, NoSQL, ETL, Luigi, Airflow, Argo, etc. Implement data storage solutions using different types of databases, such asrelational, non-relational, or cloud-based. Working collaboratively with the client and cross-functional teams to identify andaddress data-related issues and opportunities. Stay updated with the latest trends and developments in the data engineering field,such as modern data stack, big data technologies, cloud computing, etc Assist with defining the processes needed to achieve operational excellence Ensure data quality across all projects/clients Define and manage SLA's for data sets and processes throughout production Lead on the design, build and launch of new data models and pipelines Key Skills/Experience: Experience in operating as part of data engineering teams and independently. Line/Team management experience, leadership experience Experience of working with cloud infrastructure (Azure or AWS, GCP is beneficial) SQL and relational databases (e.g. MS SQL/Azure SQL, PostgreSQL) You have framework experience within either Flask, Tornado or Django, Docker Experience working with ETL pipelines is desirable e.g. Luigi, Airflow or Argo Experience with big data technologies, such as Apache Spark, Hadoop, Kafka, etc. Data acquisition and development of data sets and improving data quality Preparing data for predictive and prescriptive modelling Hands on coding experience, such as Python Reporting tools (e.g. Tableau, PowerBI, Qlik) GDPR and Government Service Standard (desirable) Passionate, motivated and enthusiastic about developing technology solutions. Experience working in an Agile development environment Data architecture experience. Competitive Salary Bonus Scheme Private Healthcare Insurance 25 Days Annual Leave + Bank Holidays Up to 10 days allocated for development training per year Enhanced Parental Leave Paid Fertility Leave (5 Days) Statutory & Contributory Pension EAP with Gym Membership Benefits Cycle to Work and Electric Vehicle schemes Flexible Working Annual Away Days/Company Socials Diversity and Inclusion As an equal opportunities employer, we are committed to creating a work environment that supports, celebrates, encourages and respects all individuals, where all processes are based on merit, competence and business needs. Encouraging high social mobility is really important to us. We foster an inclusive culture by welcoming different perspectives, enabling equitable opportunities and promoting open dialogue. This commitment is reflected in initiatives such as our gender diversity group and our focus on mental health and wellbeing. Whatever stage you are at, you will find an environment where you can thrive. Should you require further assistance or require any reasonable adjustments to be put in place to better support your application process, please do not hesitate to raise this with us. As a Disability Confident employer, we are committed to ensuring our recruitment process is accessible and inclusive, enabling all candidates to demonstrate their skills, experience and potential.
We're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people-centric and here to power good.Every day, we future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our companyand customers from what's now to what's next.We make it happen through the power of acceleration. Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us. Job Description Mandatory Skills: Observability, Resiliency, Service Management, Reliability, Performance engineering, Scalability, release management, Cloud cost management. Role Description Skills: ROLE PURPOSE Lead the Site Reliability Engineering practice, driving the transformation from reactive operations to proactive, engineering-led reliability. Own the definition and enforcement of non-functional requirements (NFRs) using FMEA-based resiliency frameworks, and champion observability, self-healing automation, automated incident management, and database operations automation. Ensure systems are resilient, performant, cost-optimised, and continuously improving. KEY RESPONSIBILITIES Define and enforce non-functional requirements (NFRs) for performance, scalability, availability, fault tolerance, and cost efficiency using FMEA-based failure analysis Design and implement self-healing automation for known failure patterns, reducing human intervention and on-call burden by 50%+ Build comprehensive observability stacks (metrics, logs, traces) with ML-driven anomaly detection and AIOps capabilities Lead automated incident management: detection, triage, escalation, remediation, and post-incident review automation Drive DB automation: automated provisioning, release management (UK focus), backup/restore, and operational request workflows for all database operations Define and track SLIs, SLOs, and error budgets across all critical services, using them to balance reliability with feature velocity Conduct chaos engineering exercises and game days to validate resiliency and uncover hidden failure modes Mentor 2 SRE Engineers, establish engineering standards, and build a culture of reliability and continuous improvement Collaborate with Platform Engineering and Cloud teams to embed reliability into infrastructure and deployment pipelines TECHNICAL SKILLS & EXPERTISE Expert-level observability: Prometheus, Grafana, ELK/OpenSearch, Jaeger/Zipkin, Datadog, or Dynatrace Strong experience with AIOps and ML-driven monitoring: PagerDuty, Moogsoft, BigPanda, or custom ML pipelines Deep knowledge of FMEA, fault tree analysis, and chaos engineering tools (Gremlin, LitmusChaos, Chaos Monkey) Database automation: strong SQL skills plus experience with automated DB provisioning, migration tools (Liquibase, Flyway), and DB release pipelines Proficiency in automation and scripting: Python, Go, Bash, with experience building self-healing runbooks Infrastructure knowledge: Kubernetes, cloud platforms (AWS/Azure/GCP), networking, and storage systems CI/CD and release engineering: Jenkins, GitLab CI, Spinnaker, ArgoCD for integrated DB and application releases Cost management: experience with FinOps principles, resource optimisation, and cloud spend analysis SOFT SKILLS & COMPETENCIES Strong leadership and mentoring ability - coaches and develops junior engineers Excellent stakeholder management and communication skills across all levels Strategic thinker who balances technical depth with business outcomes Proven ability to drive change, influence without authority, and build consensus Strong analytical and problem-solving mindset with attention to detail Ability to manage competing priorities across multiple workstreams simultaneously QUALIFICATIONS & EXPERIENCE 7+ years in SRE, DevOps, or production engineering with 3+ years in a senior or lead capacity Proven track record of improving availability, reducing MTTR, and implementing self-healing at scale Experience managing or automating database operations in enterprise environments Relevant certifications preferred: CKA, AWS DevOps Professional, Azure DevOps Expert, SRE Foundation Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience) DESIRABLE / NICE TO HAVE Published work or conference talks on SRE, observability, or chaos engineering Experience with service mesh (Istio) and distributed tracing at scale Background in financial services or regulated industry SRE practices About Us We're a global, team of innovators. Together, we harness engineering excellence and passion to co-create meaningful solutions to complex challenges. We turn organizations into data-driven leaders that can make a positive impact on their industries and society. If you believe that innovation can bring a better tomorrow closer to today, this is the place for you. Fostering innovation through diverse perspectives Hitachi is a global company operating across a wide range of industries and regions. One of the things that sets Hitachi apart is the diversity of our business and people, which drives our innovation and growth. We are committed to building an inclusive culture based on mutual respect and merit-based systems. We believe that when people feel valued, heard, and safe to express themselves, they do their best work. How we look after you We help take care of your today and tomorrow with industry-leading benefits, support, and services that look after your holistic health and wellbeing. We're also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We're always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you'll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with. We're proud to say we're an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
27/07/2026
Full time
We're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people-centric and here to power good.Every day, we future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our companyand customers from what's now to what's next.We make it happen through the power of acceleration. Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us. Job Description Mandatory Skills: Observability, Resiliency, Service Management, Reliability, Performance engineering, Scalability, release management, Cloud cost management. Role Description Skills: ROLE PURPOSE Lead the Site Reliability Engineering practice, driving the transformation from reactive operations to proactive, engineering-led reliability. Own the definition and enforcement of non-functional requirements (NFRs) using FMEA-based resiliency frameworks, and champion observability, self-healing automation, automated incident management, and database operations automation. Ensure systems are resilient, performant, cost-optimised, and continuously improving. KEY RESPONSIBILITIES Define and enforce non-functional requirements (NFRs) for performance, scalability, availability, fault tolerance, and cost efficiency using FMEA-based failure analysis Design and implement self-healing automation for known failure patterns, reducing human intervention and on-call burden by 50%+ Build comprehensive observability stacks (metrics, logs, traces) with ML-driven anomaly detection and AIOps capabilities Lead automated incident management: detection, triage, escalation, remediation, and post-incident review automation Drive DB automation: automated provisioning, release management (UK focus), backup/restore, and operational request workflows for all database operations Define and track SLIs, SLOs, and error budgets across all critical services, using them to balance reliability with feature velocity Conduct chaos engineering exercises and game days to validate resiliency and uncover hidden failure modes Mentor 2 SRE Engineers, establish engineering standards, and build a culture of reliability and continuous improvement Collaborate with Platform Engineering and Cloud teams to embed reliability into infrastructure and deployment pipelines TECHNICAL SKILLS & EXPERTISE Expert-level observability: Prometheus, Grafana, ELK/OpenSearch, Jaeger/Zipkin, Datadog, or Dynatrace Strong experience with AIOps and ML-driven monitoring: PagerDuty, Moogsoft, BigPanda, or custom ML pipelines Deep knowledge of FMEA, fault tree analysis, and chaos engineering tools (Gremlin, LitmusChaos, Chaos Monkey) Database automation: strong SQL skills plus experience with automated DB provisioning, migration tools (Liquibase, Flyway), and DB release pipelines Proficiency in automation and scripting: Python, Go, Bash, with experience building self-healing runbooks Infrastructure knowledge: Kubernetes, cloud platforms (AWS/Azure/GCP), networking, and storage systems CI/CD and release engineering: Jenkins, GitLab CI, Spinnaker, ArgoCD for integrated DB and application releases Cost management: experience with FinOps principles, resource optimisation, and cloud spend analysis SOFT SKILLS & COMPETENCIES Strong leadership and mentoring ability - coaches and develops junior engineers Excellent stakeholder management and communication skills across all levels Strategic thinker who balances technical depth with business outcomes Proven ability to drive change, influence without authority, and build consensus Strong analytical and problem-solving mindset with attention to detail Ability to manage competing priorities across multiple workstreams simultaneously QUALIFICATIONS & EXPERIENCE 7+ years in SRE, DevOps, or production engineering with 3+ years in a senior or lead capacity Proven track record of improving availability, reducing MTTR, and implementing self-healing at scale Experience managing or automating database operations in enterprise environments Relevant certifications preferred: CKA, AWS DevOps Professional, Azure DevOps Expert, SRE Foundation Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience) DESIRABLE / NICE TO HAVE Published work or conference talks on SRE, observability, or chaos engineering Experience with service mesh (Istio) and distributed tracing at scale Background in financial services or regulated industry SRE practices About Us We're a global, team of innovators. Together, we harness engineering excellence and passion to co-create meaningful solutions to complex challenges. We turn organizations into data-driven leaders that can make a positive impact on their industries and society. If you believe that innovation can bring a better tomorrow closer to today, this is the place for you. Fostering innovation through diverse perspectives Hitachi is a global company operating across a wide range of industries and regions. One of the things that sets Hitachi apart is the diversity of our business and people, which drives our innovation and growth. We are committed to building an inclusive culture based on mutual respect and merit-based systems. We believe that when people feel valued, heard, and safe to express themselves, they do their best work. How we look after you We help take care of your today and tomorrow with industry-leading benefits, support, and services that look after your holistic health and wellbeing. We're also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We're always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you'll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with. We're proud to say we're an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
Job Title Senior Observability Engineer Job Description Senior Observability Engineer Location London Employment type Permanent, Full Time Reporting into Senior Engineering Manager - SRE and Observability About IG Group IG Group (LSE: IGG) is a leading global fintech company, established in 1974 and headquartered in London. As a constituent of the FTSE 100, IG Group provides dynamic online trading platforms and a robust educational ecosystem, empowering ambitious individuals worldwide in their pursuit of financial freedom. With operations spanning 18 countries across Europe, Africa, Asia Pacific, the Middle East, and North America, IG Group offers clients access to approximately 19,000 financial markets, including shares, forex, indices, and commodities. IG Group is undergoing a significant transformation, led by CEO Breon Corcoran, appointed in early 2024. Our focus is on three areas: delivering quality products to better meet customer needs, embedding a high-performance culture across the organisation, and being more efficient and scalable through digitisation - all in service of growing our user base and revenue on a sustainable basis. About the role IG Group's systems move billions of dollars every day - and our clients expect them to be fast, reliable, and transparent. As an Observability Engineer, you will own the platforms and practices that give IG's engineering teams deep, real-time insight into how those systems behave. This is a high, hands-on role at the centre of our reliability engineering agenda: building on Honeycomb and OpenTelemetry, driving instrumentation across a globally distributed microservices estate, and partnering directly with development teams to turn telemetry data into better, faster software. About the team This role sits within the Observability team, part of IG's broader SRE and Platform Engineering function. The team is responsible for the tools, platforms, and standards that enable engineering teams across IG to understand and improve system behaviour at scale. You will report into the Senior Engineering Manager for SRE and Observability and work as an individual contributor, partnering closely with development squads, platform engineers, and incident response teams. The Bengaluru team is deeply integrated into IG's global engineering community, with real ownership and scope to shape how observability is practised across the group. Key responsibilities Platform ownership - Build, maintain, and evolve IG's Honeycomb-centric observability platform, ensuring it is reliable, scalable, and fit for a complex, globally distributed trading environment. Define platform standards, data models, and integration patterns for telemetry collection, storage, and querying across the estate. Instrumentation and telemetry - Drive OpenTelemetry adoption across engineering teams, providing hands-on guidance and reusable instrumentation patterns for services built in Java, Python, and C++. Partner with development teams to improve telemetry coverage, ensuring meaningful traces, metrics, and logs are in place across critical services and user journeys. SLOs - Apply a strong understanding of SLOs, burn rates, and alert triggers to help service owners define meaningful reliability targets and translate them into actionable observability signals. Work closely with SRE teams to drive SLO adoption across the organisation, providing guidance and practical support to help engineering teams embed reliability targets into their day-to-day ways of working. Incident response - Join the support rota and incident response, using observability tooling to accelerate diagnosis and reduce mean time to resolution. Lead post-incident reviews that produce actionable improvements to both systems and observability coverage. Enablement and community - Mentor and upskill engineers across IG on observability principles and practices, raising the bar for how teams instrument, monitor, and debug their services. Develop training materials, runbooks, and best- practice guides that scale observability knowledge across the engineering organisation. Role requirements Proven hands-on experience with Honeycomb or a similar observability tool such as Grafana, including dataset design, query building, and using it as a primary tool for production debugging and reliability analysis. Strong practical experience implementing OpenTelemetry instrumentation in one or more of Java, Python, or C++, including custom collectors, exporters, and sampling strategies. Experience working in complex, distributed microservices environments with high transaction volumes or strict reliability requirements. Strong communication and collaboration skills - able to work effectively with development teams and influence engineering practice without direct authority. Practical experience with Terraform for managing observability infrastructure as code, including provisioning and maintaining platform components in a cloud environment. Ability and willingness to cover UK working hours to support collaboration with IG's London-based engineering teams. 5-8 years of relevant experience in observability, SRE, or platform engineering roles. Desirable Experience with cloud platforms such as AWS or GCP, particularly in the context of observability and infrastructure monitoring. Familiarity with other observability tooling such as Grafana, Prometheus, or Splunk, and experience migrating or consolidating observability stacks. Exposure to fintech, financial services, or other regulated industry environments. Experience contributing to or maintaining open-source observability projects or OTel instrumentation libraries. The Perks Your growth fuels our success! Thrive with tailored development programs, mentoring opportunities with leaders, and clear career progression. Expand your network through committees, sports and social clubs. Enjoy extra time off for volunteering and community work. Competitive salary Flexible Benefits Package on top of your salary (12%) Private medical cover for you and your family Life insurance Contribution to gym memberships 25 Days holiday, with 1 additional day off to celebrate your Birthday & 2 additional days off a year for voluntary work (28 in total The option to buy or sell holiday days. Unlimited access to the LinkedIn Learning Platform A comprehensive global and local onboarding process Employee-led LGBTQ+, Women's, Black and Parents & Carers networks with an annual budget for organising events & projects that foster an open, diverse and inclusive culture Enhanced primary (maternity), secondary (paternity), and shared parental pay and leave, as well as a range of support and benefits for parents Option to participate and create ESG initiatives based on IG Brighter Future Fund Number of openings 1 Looking for a career at a company that will support you, challenge you and help you grow? IG Group can provide that. IG Group is a FTSE 100 fintech operating across five continents, serving over 1.3m customers and handling billions of dollars in transactions - built on scale, trust, and proof. We didn't pivot to innovation; it's how we've always operated. What that means for the people who work here is real: genuinely complex problems to solve, the technology and resources to tackle them properly, and the kind of scope that's rare in established businesses. The bar is high - bring a curious and forward-thinking mindset and we'll give you the platform to define what comes next. Join us at IG - the future gets built here.
26/07/2026
Full time
Job Title Senior Observability Engineer Job Description Senior Observability Engineer Location London Employment type Permanent, Full Time Reporting into Senior Engineering Manager - SRE and Observability About IG Group IG Group (LSE: IGG) is a leading global fintech company, established in 1974 and headquartered in London. As a constituent of the FTSE 100, IG Group provides dynamic online trading platforms and a robust educational ecosystem, empowering ambitious individuals worldwide in their pursuit of financial freedom. With operations spanning 18 countries across Europe, Africa, Asia Pacific, the Middle East, and North America, IG Group offers clients access to approximately 19,000 financial markets, including shares, forex, indices, and commodities. IG Group is undergoing a significant transformation, led by CEO Breon Corcoran, appointed in early 2024. Our focus is on three areas: delivering quality products to better meet customer needs, embedding a high-performance culture across the organisation, and being more efficient and scalable through digitisation - all in service of growing our user base and revenue on a sustainable basis. About the role IG Group's systems move billions of dollars every day - and our clients expect them to be fast, reliable, and transparent. As an Observability Engineer, you will own the platforms and practices that give IG's engineering teams deep, real-time insight into how those systems behave. This is a high, hands-on role at the centre of our reliability engineering agenda: building on Honeycomb and OpenTelemetry, driving instrumentation across a globally distributed microservices estate, and partnering directly with development teams to turn telemetry data into better, faster software. About the team This role sits within the Observability team, part of IG's broader SRE and Platform Engineering function. The team is responsible for the tools, platforms, and standards that enable engineering teams across IG to understand and improve system behaviour at scale. You will report into the Senior Engineering Manager for SRE and Observability and work as an individual contributor, partnering closely with development squads, platform engineers, and incident response teams. The Bengaluru team is deeply integrated into IG's global engineering community, with real ownership and scope to shape how observability is practised across the group. Key responsibilities Platform ownership - Build, maintain, and evolve IG's Honeycomb-centric observability platform, ensuring it is reliable, scalable, and fit for a complex, globally distributed trading environment. Define platform standards, data models, and integration patterns for telemetry collection, storage, and querying across the estate. Instrumentation and telemetry - Drive OpenTelemetry adoption across engineering teams, providing hands-on guidance and reusable instrumentation patterns for services built in Java, Python, and C++. Partner with development teams to improve telemetry coverage, ensuring meaningful traces, metrics, and logs are in place across critical services and user journeys. SLOs - Apply a strong understanding of SLOs, burn rates, and alert triggers to help service owners define meaningful reliability targets and translate them into actionable observability signals. Work closely with SRE teams to drive SLO adoption across the organisation, providing guidance and practical support to help engineering teams embed reliability targets into their day-to-day ways of working. Incident response - Join the support rota and incident response, using observability tooling to accelerate diagnosis and reduce mean time to resolution. Lead post-incident reviews that produce actionable improvements to both systems and observability coverage. Enablement and community - Mentor and upskill engineers across IG on observability principles and practices, raising the bar for how teams instrument, monitor, and debug their services. Develop training materials, runbooks, and best- practice guides that scale observability knowledge across the engineering organisation. Role requirements Proven hands-on experience with Honeycomb or a similar observability tool such as Grafana, including dataset design, query building, and using it as a primary tool for production debugging and reliability analysis. Strong practical experience implementing OpenTelemetry instrumentation in one or more of Java, Python, or C++, including custom collectors, exporters, and sampling strategies. Experience working in complex, distributed microservices environments with high transaction volumes or strict reliability requirements. Strong communication and collaboration skills - able to work effectively with development teams and influence engineering practice without direct authority. Practical experience with Terraform for managing observability infrastructure as code, including provisioning and maintaining platform components in a cloud environment. Ability and willingness to cover UK working hours to support collaboration with IG's London-based engineering teams. 5-8 years of relevant experience in observability, SRE, or platform engineering roles. Desirable Experience with cloud platforms such as AWS or GCP, particularly in the context of observability and infrastructure monitoring. Familiarity with other observability tooling such as Grafana, Prometheus, or Splunk, and experience migrating or consolidating observability stacks. Exposure to fintech, financial services, or other regulated industry environments. Experience contributing to or maintaining open-source observability projects or OTel instrumentation libraries. The Perks Your growth fuels our success! Thrive with tailored development programs, mentoring opportunities with leaders, and clear career progression. Expand your network through committees, sports and social clubs. Enjoy extra time off for volunteering and community work. Competitive salary Flexible Benefits Package on top of your salary (12%) Private medical cover for you and your family Life insurance Contribution to gym memberships 25 Days holiday, with 1 additional day off to celebrate your Birthday & 2 additional days off a year for voluntary work (28 in total The option to buy or sell holiday days. Unlimited access to the LinkedIn Learning Platform A comprehensive global and local onboarding process Employee-led LGBTQ+, Women's, Black and Parents & Carers networks with an annual budget for organising events & projects that foster an open, diverse and inclusive culture Enhanced primary (maternity), secondary (paternity), and shared parental pay and leave, as well as a range of support and benefits for parents Option to participate and create ESG initiatives based on IG Brighter Future Fund Number of openings 1 Looking for a career at a company that will support you, challenge you and help you grow? IG Group can provide that. IG Group is a FTSE 100 fintech operating across five continents, serving over 1.3m customers and handling billions of dollars in transactions - built on scale, trust, and proof. We didn't pivot to innovation; it's how we've always operated. What that means for the people who work here is real: genuinely complex problems to solve, the technology and resources to tackle them properly, and the kind of scope that's rare in established businesses. The bar is high - bring a curious and forward-thinking mindset and we'll give you the platform to define what comes next. Join us at IG - the future gets built here.
FunctionCloud & Data EngineeringOur CompanyWe're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people-centric and here to power good. Every day, we future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our company and customers from what's now to what's next. We make it happen through the power of acceleration.Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us.Job descriptionMandatory Skills:Observability, Resiliency, Service Management, Reliability, Performance engineering, Scalability, release management, Cloud cost management.Role Description Skills:ROLE PURPOSELead the Site Reliability Engineering practice, driving the transformation from reactive operations to proactive, engineering-led reliability. Own the definition and enforcement of non-functional requirements (NFRs) using FMEA-based resiliency frameworks, and champion observability, self-healing automation, automated incident management, and database operations automation. Ensure systems are resilient, performant, cost-optimised, and continuously improving.KEY RESPONSIBILITIESDefine and enforce non-functional requirements (NFRs) for performance, scalability, availability, fault tolerance, and cost efficiency using FMEA-based failure analysisDesign and implement self-healing automation for known failure patterns, reducing human intervention and on-call burden by 50%+Build comprehensive observability stacks (metrics, logs, traces) with ML-driven anomaly detection and AIOps capabilitiesLead automated incident management: detection, triage, escalation, remediation, and post-incident review automationDrive DB automation: automated provisioning, release management (UK focus), backup/restore, and operational request workflows for all database operationsDefine and track SLIs, SLOs, and error budgets across all critical services, using them to balance reliability with feature velocityConduct chaos engineering exercises and game days to validate resiliency and uncover hidden failure modesMentor 2 SRE Engineers, establish engineering standards, and build a culture of reliability and continuous improvementCollaborate with Platform Engineering and Cloud teams to embed reliability into infrastructure and deployment pipelinesTECHNICAL SKILLS & EXPERTISEExpert-level observability: Prometheus, Grafana, ELK/OpenSearch, Jaeger/Zipkin, Datadog, or DynatraceStrong experience with AIOps and ML-driven monitoring: PagerDuty, Moogsoft, BigPanda, or custom ML pipelinesDeep knowledge of FMEA, fault tree analysis, and chaos engineering tools (Gremlin, LitmusChaos, Chaos Monkey)Database automation: strong SQL skills plus experience with automated DB provisioning, migration tools (Liquibase, Flyway), and DB release pipelinesProficiency in automation and scripting: Python, Go, Bash, with experience building self-healing runbooksInfrastructure knowledge: Kubernetes, cloud platforms (AWS/Azure/GCP), networking, and storage systemsCI/CD and release engineering: Jenkins, GitLab CI, Spinnaker, ArgoCD for integrated DB and application releasesCost management: experience with FinOps principles, resource optimisation, and cloud spend analysisSOFT SKILLS & COMPETENCIESStrong leadership and mentoring ability - coaches and develops junior engineersExcellent stakeholder management and communication skills across all levelsStrategic thinker who balances technical depth with business outcomesProven ability to drive change, influence without authority, and build consensusStrong analytical and problem-solving mindset with attention to detailAbility to manage competing priorities across multiple workstreams simultaneouslyQUALIFICATIONS & EXPERIENCE7+ years in SRE, DevOps, or production engineering with 3+ years in a senior or lead capacityProven track record of improving availability, reducing MTTR, and implementing self-healing at scaleExperience managing or automating database operations in enterprise environmentsRelevant certifications preferred: CKA, AWS DevOps Professional, Azure DevOps Expert, SRE FoundationBachelor's degree in Computer Science, Engineering, or related field (or equivalent experience)DESIRABLE / NICE TO HAVEPublished work or conference talks on SRE, observability, or chaos engineeringExperience with service mesh (Istio) and distributed tracing at scaleBackground in financial services or regulated industry SRE practicesAbout usWe're a global, team of innovators. Together, we harness engineering excellence and passion to co-create meaningful solutions to complex challenges. We turn organizations into data-driven leaders that can make a positive impact on their industries and society. If you believe that innovation can bring a better tomorrow closer to today, this is the place for you.Fostering innovation through diverse perspectivesHitachi is a global company operating across a wide range of industries and regions. One of the things that sets Hitachi apart is the diversity of our business and people, which drives our innovation and growth.We are committed to building an inclusive culture based on mutual respect and merit-based systems. We believe that when people feel valued, heard, and safe to express themselves, they do their best work.How we look after youWe help take care of your today and tomorrow with industry-leading benefits, support, and services that look after your holistic health and wellbeing. We're also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We're always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you'll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with.We're proud to say we're an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
26/07/2026
Full time
FunctionCloud & Data EngineeringOur CompanyWe're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people-centric and here to power good. Every day, we future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our company and customers from what's now to what's next. We make it happen through the power of acceleration.Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us.Job descriptionMandatory Skills:Observability, Resiliency, Service Management, Reliability, Performance engineering, Scalability, release management, Cloud cost management.Role Description Skills:ROLE PURPOSELead the Site Reliability Engineering practice, driving the transformation from reactive operations to proactive, engineering-led reliability. Own the definition and enforcement of non-functional requirements (NFRs) using FMEA-based resiliency frameworks, and champion observability, self-healing automation, automated incident management, and database operations automation. Ensure systems are resilient, performant, cost-optimised, and continuously improving.KEY RESPONSIBILITIESDefine and enforce non-functional requirements (NFRs) for performance, scalability, availability, fault tolerance, and cost efficiency using FMEA-based failure analysisDesign and implement self-healing automation for known failure patterns, reducing human intervention and on-call burden by 50%+Build comprehensive observability stacks (metrics, logs, traces) with ML-driven anomaly detection and AIOps capabilitiesLead automated incident management: detection, triage, escalation, remediation, and post-incident review automationDrive DB automation: automated provisioning, release management (UK focus), backup/restore, and operational request workflows for all database operationsDefine and track SLIs, SLOs, and error budgets across all critical services, using them to balance reliability with feature velocityConduct chaos engineering exercises and game days to validate resiliency and uncover hidden failure modesMentor 2 SRE Engineers, establish engineering standards, and build a culture of reliability and continuous improvementCollaborate with Platform Engineering and Cloud teams to embed reliability into infrastructure and deployment pipelinesTECHNICAL SKILLS & EXPERTISEExpert-level observability: Prometheus, Grafana, ELK/OpenSearch, Jaeger/Zipkin, Datadog, or DynatraceStrong experience with AIOps and ML-driven monitoring: PagerDuty, Moogsoft, BigPanda, or custom ML pipelinesDeep knowledge of FMEA, fault tree analysis, and chaos engineering tools (Gremlin, LitmusChaos, Chaos Monkey)Database automation: strong SQL skills plus experience with automated DB provisioning, migration tools (Liquibase, Flyway), and DB release pipelinesProficiency in automation and scripting: Python, Go, Bash, with experience building self-healing runbooksInfrastructure knowledge: Kubernetes, cloud platforms (AWS/Azure/GCP), networking, and storage systemsCI/CD and release engineering: Jenkins, GitLab CI, Spinnaker, ArgoCD for integrated DB and application releasesCost management: experience with FinOps principles, resource optimisation, and cloud spend analysisSOFT SKILLS & COMPETENCIESStrong leadership and mentoring ability - coaches and develops junior engineersExcellent stakeholder management and communication skills across all levelsStrategic thinker who balances technical depth with business outcomesProven ability to drive change, influence without authority, and build consensusStrong analytical and problem-solving mindset with attention to detailAbility to manage competing priorities across multiple workstreams simultaneouslyQUALIFICATIONS & EXPERIENCE7+ years in SRE, DevOps, or production engineering with 3+ years in a senior or lead capacityProven track record of improving availability, reducing MTTR, and implementing self-healing at scaleExperience managing or automating database operations in enterprise environmentsRelevant certifications preferred: CKA, AWS DevOps Professional, Azure DevOps Expert, SRE FoundationBachelor's degree in Computer Science, Engineering, or related field (or equivalent experience)DESIRABLE / NICE TO HAVEPublished work or conference talks on SRE, observability, or chaos engineeringExperience with service mesh (Istio) and distributed tracing at scaleBackground in financial services or regulated industry SRE practicesAbout usWe're a global, team of innovators. Together, we harness engineering excellence and passion to co-create meaningful solutions to complex challenges. We turn organizations into data-driven leaders that can make a positive impact on their industries and society. If you believe that innovation can bring a better tomorrow closer to today, this is the place for you.Fostering innovation through diverse perspectivesHitachi is a global company operating across a wide range of industries and regions. One of the things that sets Hitachi apart is the diversity of our business and people, which drives our innovation and growth.We are committed to building an inclusive culture based on mutual respect and merit-based systems. We believe that when people feel valued, heard, and safe to express themselves, they do their best work.How we look after youWe help take care of your today and tomorrow with industry-leading benefits, support, and services that look after your holistic health and wellbeing. We're also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We're always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you'll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with.We're proud to say we're an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
The position works closely with IAM architects, business stakeholders and technology partners. In this role, you will: Build and maintain cloud-native platform services supporting the IDAM 2.0 programme. Design, deploy and manage Kubernetes clusters and supporting platform components. Develop Infrastructure as Code (IaC) for repeatable, automated deployments. Implement and maintain CI/CD pipelines and GitOps deployment workflows. Manage cloud networking, connectivity and platform security. Implement platform observability including logging, monitoring, metrics and distributed tracing. Automate platform provisioning, configuration management and operational tasks. Support deployment and operation of identity platform components and supporting services. Implement secure secrets management and certificate lifecycle automation. Configure service mesh technologies and secure service-to-service communication. Manage ingress, API gateways and network routing components. Implement platform resilience, backup, disaster recovery and high availability capabilities. Support workload identity and platform authentication mechanisms. Ensure platform compliance with enterprise security, governance and regulatory requirements. Optimise platform performance, scalability and cost efficiency. Develop operational tooling, scripts and self-service capabilities. Support incident management, troubleshooting and root cause analysis. Collaborate with Security, Architecture and Engineering teams to deliver secure platform services. Contribute to platform standards, operational documentation and runbooks. Support environment management across development, test, pilot and production environments. Drive continuous improvement through automation and platform engineering best practices. To be successful in this role, you should meet the following requirements: Key Skills & Experience Essential Skills Cloud Platform Engineering (GCP) Kubernetes Administration Infrastructure as Code (Terraform or equivalent) CI/CD Pipelines GitOps Linux Administration Networking and Load Balancing Service Mesh Technologies (Istio/Envoy) Container Platforms Secrets Management PKI and Certificate Management Workload Identity Observability (Logging, Monitoring and Tracing) Scripting and Automation DevSecOps Site Reliability Engineering (SRE) Performance and Capacity Management Operational Support Global Deployment Strategies Positive can-do attitude adept at solutionising in large, complex organisations
26/07/2026
Full time
The position works closely with IAM architects, business stakeholders and technology partners. In this role, you will: Build and maintain cloud-native platform services supporting the IDAM 2.0 programme. Design, deploy and manage Kubernetes clusters and supporting platform components. Develop Infrastructure as Code (IaC) for repeatable, automated deployments. Implement and maintain CI/CD pipelines and GitOps deployment workflows. Manage cloud networking, connectivity and platform security. Implement platform observability including logging, monitoring, metrics and distributed tracing. Automate platform provisioning, configuration management and operational tasks. Support deployment and operation of identity platform components and supporting services. Implement secure secrets management and certificate lifecycle automation. Configure service mesh technologies and secure service-to-service communication. Manage ingress, API gateways and network routing components. Implement platform resilience, backup, disaster recovery and high availability capabilities. Support workload identity and platform authentication mechanisms. Ensure platform compliance with enterprise security, governance and regulatory requirements. Optimise platform performance, scalability and cost efficiency. Develop operational tooling, scripts and self-service capabilities. Support incident management, troubleshooting and root cause analysis. Collaborate with Security, Architecture and Engineering teams to deliver secure platform services. Contribute to platform standards, operational documentation and runbooks. Support environment management across development, test, pilot and production environments. Drive continuous improvement through automation and platform engineering best practices. To be successful in this role, you should meet the following requirements: Key Skills & Experience Essential Skills Cloud Platform Engineering (GCP) Kubernetes Administration Infrastructure as Code (Terraform or equivalent) CI/CD Pipelines GitOps Linux Administration Networking and Load Balancing Service Mesh Technologies (Istio/Envoy) Container Platforms Secrets Management PKI and Certificate Management Workload Identity Observability (Logging, Monitoring and Tracing) Scripting and Automation DevSecOps Site Reliability Engineering (SRE) Performance and Capacity Management Operational Support Global Deployment Strategies Positive can-do attitude adept at solutionising in large, complex organisations
Senior Observability EngineerApplylocations: City of London - United Kingdomtime type: Full timeposted on: Posted Todayjob requisition id: R\_17538 Job Title Senior Observability Engineer Job Description Senior Observability Engineer Location: London Employment type: Permanent, Full Time Reporting into: Senior Engineering Manager - SRE and Observability About IG Group IG Group (LSE: IGG) is a leading global fintech company, established in 1974 and headquartered in London. As a constituent of the FTSE 100, IG Group provides dynamic online trading platforms and a robust educational ecosystem, empowering ambitious individuals worldwide in their pursuit of financial freedom. With operations spanning 18 countries across Europe, Africa, Asia-Pacific, the Middle East, and North America, IG Group offers clients access to approximately 19,000 financial markets, including shares, forex, indices, and commodities.IG Group is undergoing a significant transformation, led by CEO Breon Corcoran, appointed in early 2024. Our focus is on three areas: delivering quality products to better meet customer needs, embedding a high-performance culture across the organisation, and being more efficient and scalable through digitisation - all in service of growing our user base and revenue on a sustainable basis. About the role IG Group's systems move billions of dollars every day - and our clients expect them to be fast, reliable, and transparent. As an Observability Engineer, you will own the platforms and practices that give IG's engineering teams deep, real-time insight into how those systems behave. This is a high-impact, hands-on role at the centre of our reliability engineering agenda: building on Honeycomb and OpenTelemetry, driving instrumentation across a globally distributed microservices estate, and partnering directly with development teams to turn telemetry data into better, faster software. About the team This role sits within the Observability team, part of IG's broader SRE and Platform Engineering function. The team is responsible for the tools, platforms, and standards that enable engineering teams across IG to understand and improve system behaviour at scale. You will report into the Senior Engineering Manager for SRE and Observability and work as an individual contributor, partnering closely with development squads, platform engineers, and incident response teams. The Bengaluru team is deeply integrated into IG's global engineering community, with real ownership and scope to shape how observability is practised across the group. Key responsibilities Platform ownership Build, maintain, and evolve IG's Honeycomb-centric observability platform, ensuring it is reliable, scalable, and fit for a complex, globally distributed trading environment. Define platform standards, data models, and integration patterns for telemetry collection, storage, and querying across the estate. Instrumentation and telemetry Drive OpenTelemetry adoption across engineering teams, providing hands-on guidance and reusable instrumentation patterns for services built in Java, Python, and C++. Partner with development teams to improve telemetry coverage, ensuring meaningful traces, metrics, and logs are in place across critical services and user journeys. SLOs Apply a strong understanding of SLOs, burn rates, and alert triggers to help service owners define meaningful reliability targets and translate them into actionable observability signals. Work closely with SRE teams to drive SLO adoption across the organisation, providing guidance and practical support to help engineering teams embed reliability targets into their day-to-day ways of working. Incident response Join the support rota and incident response, using observability tooling to accelerate diagnosis and reduce mean time to resolution. Lead post-incident reviews that produce actionable improvements to both systems and observability coverage. Enablement and community Mentor and upskill engineers across IG on observability principles and practices, raising the bar for how teams instrument, monitor, and debug their services. Develop training materials, runbooks, and best-practice guides that scale observability knowledge across the engineering organisation. Role requirements Proven hands-on experience with Honeycomb or a similar observability tool such as Grafana, including dataset design, query building, and using it as a primary tool for production debugging and reliability analysis. Strong practical experience implementing OpenTelemetry instrumentation in one or more of Java, Python, or C++, including custom collectors, exporters, and sampling strategies. Experience working in complex, distributed microservices environments with high transaction volumes or strict reliability requirements. Strong communication and collaboration skills - able to work effectively with development teams and influence engineering practice without direct authority. Practical experience with Terraform for managing observability infrastructure as code, including provisioning and maintaining platform components in a cloud environment. Ability and willingness to cover UK working hours to support collaboration with IG's London-based engineering teams. 5-8 years of relevant experience in observability, SRE, or platform engineering roles. Desirable Experience with cloud platforms such as AWS or GCP, particularly in the context of observability and infrastructure monitoring. Familiarity with other observability tooling such as Grafana, Prometheus, or Splunk, and experience migrating or consolidating observability stacks. Exposure to fintech, financial services, or other regulated industry environments. Experience contributing to or maintaining open-source observability projects or OTel instrumentation libraries.# The Perks Your growth fuels our success! Thrive with tailored development programs, mentoring opportunities with leaders, and clear career progression. Expand your network through committees, sports and social clubs. Enjoy extra time off for volunteering and community work. Competitive salary Flexible Benefits Package on top of your salary (12%) Private medical cover for you and your family Life insurance Contribution to gym memberships 25 Days holiday, with 1 additional day off to celebrate your Birthday & 2 additional days off a year for voluntary work (28 in total The option to buy or sell holiday days. Unlimited access to the LinkedIn Learning Platform A comprehensive global and local onboarding process Employee-led LGBTQ+, Women's, Black and Parents & Carers networks with an annual budget for organising events & projects that foster an open, diverse and inclusive culture Enhanced primary (maternity), secondary (paternity), and shared parental pay and leave, as well as a range of support and benefits for parents Option to participate and create ESG initiatives based on IG Brighter Future Fund
26/07/2026
Full time
Senior Observability EngineerApplylocations: City of London - United Kingdomtime type: Full timeposted on: Posted Todayjob requisition id: R\_17538 Job Title Senior Observability Engineer Job Description Senior Observability Engineer Location: London Employment type: Permanent, Full Time Reporting into: Senior Engineering Manager - SRE and Observability About IG Group IG Group (LSE: IGG) is a leading global fintech company, established in 1974 and headquartered in London. As a constituent of the FTSE 100, IG Group provides dynamic online trading platforms and a robust educational ecosystem, empowering ambitious individuals worldwide in their pursuit of financial freedom. With operations spanning 18 countries across Europe, Africa, Asia-Pacific, the Middle East, and North America, IG Group offers clients access to approximately 19,000 financial markets, including shares, forex, indices, and commodities.IG Group is undergoing a significant transformation, led by CEO Breon Corcoran, appointed in early 2024. Our focus is on three areas: delivering quality products to better meet customer needs, embedding a high-performance culture across the organisation, and being more efficient and scalable through digitisation - all in service of growing our user base and revenue on a sustainable basis. About the role IG Group's systems move billions of dollars every day - and our clients expect them to be fast, reliable, and transparent. As an Observability Engineer, you will own the platforms and practices that give IG's engineering teams deep, real-time insight into how those systems behave. This is a high-impact, hands-on role at the centre of our reliability engineering agenda: building on Honeycomb and OpenTelemetry, driving instrumentation across a globally distributed microservices estate, and partnering directly with development teams to turn telemetry data into better, faster software. About the team This role sits within the Observability team, part of IG's broader SRE and Platform Engineering function. The team is responsible for the tools, platforms, and standards that enable engineering teams across IG to understand and improve system behaviour at scale. You will report into the Senior Engineering Manager for SRE and Observability and work as an individual contributor, partnering closely with development squads, platform engineers, and incident response teams. The Bengaluru team is deeply integrated into IG's global engineering community, with real ownership and scope to shape how observability is practised across the group. Key responsibilities Platform ownership Build, maintain, and evolve IG's Honeycomb-centric observability platform, ensuring it is reliable, scalable, and fit for a complex, globally distributed trading environment. Define platform standards, data models, and integration patterns for telemetry collection, storage, and querying across the estate. Instrumentation and telemetry Drive OpenTelemetry adoption across engineering teams, providing hands-on guidance and reusable instrumentation patterns for services built in Java, Python, and C++. Partner with development teams to improve telemetry coverage, ensuring meaningful traces, metrics, and logs are in place across critical services and user journeys. SLOs Apply a strong understanding of SLOs, burn rates, and alert triggers to help service owners define meaningful reliability targets and translate them into actionable observability signals. Work closely with SRE teams to drive SLO adoption across the organisation, providing guidance and practical support to help engineering teams embed reliability targets into their day-to-day ways of working. Incident response Join the support rota and incident response, using observability tooling to accelerate diagnosis and reduce mean time to resolution. Lead post-incident reviews that produce actionable improvements to both systems and observability coverage. Enablement and community Mentor and upskill engineers across IG on observability principles and practices, raising the bar for how teams instrument, monitor, and debug their services. Develop training materials, runbooks, and best-practice guides that scale observability knowledge across the engineering organisation. Role requirements Proven hands-on experience with Honeycomb or a similar observability tool such as Grafana, including dataset design, query building, and using it as a primary tool for production debugging and reliability analysis. Strong practical experience implementing OpenTelemetry instrumentation in one or more of Java, Python, or C++, including custom collectors, exporters, and sampling strategies. Experience working in complex, distributed microservices environments with high transaction volumes or strict reliability requirements. Strong communication and collaboration skills - able to work effectively with development teams and influence engineering practice without direct authority. Practical experience with Terraform for managing observability infrastructure as code, including provisioning and maintaining platform components in a cloud environment. Ability and willingness to cover UK working hours to support collaboration with IG's London-based engineering teams. 5-8 years of relevant experience in observability, SRE, or platform engineering roles. Desirable Experience with cloud platforms such as AWS or GCP, particularly in the context of observability and infrastructure monitoring. Familiarity with other observability tooling such as Grafana, Prometheus, or Splunk, and experience migrating or consolidating observability stacks. Exposure to fintech, financial services, or other regulated industry environments. Experience contributing to or maintaining open-source observability projects or OTel instrumentation libraries.# The Perks Your growth fuels our success! Thrive with tailored development programs, mentoring opportunities with leaders, and clear career progression. Expand your network through committees, sports and social clubs. Enjoy extra time off for volunteering and community work. Competitive salary Flexible Benefits Package on top of your salary (12%) Private medical cover for you and your family Life insurance Contribution to gym memberships 25 Days holiday, with 1 additional day off to celebrate your Birthday & 2 additional days off a year for voluntary work (28 in total The option to buy or sell holiday days. Unlimited access to the LinkedIn Learning Platform A comprehensive global and local onboarding process Employee-led LGBTQ+, Women's, Black and Parents & Carers networks with an annual budget for organising events & projects that foster an open, diverse and inclusive culture Enhanced primary (maternity), secondary (paternity), and shared parental pay and leave, as well as a range of support and benefits for parents Option to participate and create ESG initiatives based on IG Brighter Future Fund
As a global leader in cybersecurity, CrowdStrike protects the people, processes and technologies that drive modern organizations. Since 2011, our mission hasn't changed - we're here to stop breaches, and we've redefined modern security with the world's most advanced AI-native platform. We work on large scale distributed systems, processing almost 3 trillion events per day and this traffic is growing daily. Our customers span all industries, and they count on CrowdStrike to keep their businesses running, their communities safe and their lives moving forward. We're proud to work for a mission-driven company leveraging AI to transform the way we work. CrowdStrikers drive their careers through flexibility and autonomy while also being expected to contribute to a culture of responsible AI adoption, experimentation, and innovation. We use an AI-first mindset as a force multiplier to proactively and continuously accelerate execution, build expertise, uncover insights, and solve complex problems. We're always looking to add talented CrowdStrikers to the team who have limitless passion, a relentless focus on innovation and a fanatical commitment to our customers, our community and each other. Ready to join a mission that matters? The future of cybersecurity starts with you. About the Role We are seeking a talented Senior Software Engineer to join our Real-Time Streaming Platform team. Our mission is to provide enterprise-grade streaming processing capabilities that empower teams across CrowdStrike to build real-time data pipelines, perform complex event processing, and derive actionable insights from massive data streams. You'll be building the foundational infrastructure that enables our security products to detect threats, analyze patterns, and respond to incidents in real-time. You'll be designing and developing highly scalable streaming services using technologies like Apache Kafka, Apache Flink, Apache Storm and Apache Spark in Java, Python and Go. You'll design and implement data pipelines that process trillions of events daily, ensuring low latency, high throughput, and fault tolerance. Your work will directly impact how CrowdStrike ingests, processes, and analyzes security telemetry at petabyte scale across our global customer base. Have no fear if you haven't worked with all these technologies - we value strong distributed systems fundamentals, the ability to write production-quality code, and the passion for solving complex engineering challenges. We'll support your growth in streaming technologies and expect you'll be comfortable collaborating with teams distributed across various geographies and time zones. Sounds like a challenge? It is! If that's what you're after, we'd love to hear from you. What You'll Do Design, develop, and maintain streaming microservices in Java and Go that provide real-time data processing capabilities to internal teams Build and optimize Apache Flink and Apache Spark based data pipelines running on Kubernetes infrastructure Design event-driven systems leveraging Kafka for high-throughput, low-latency message streaming Take end-to-end ownership of technical initiatives, from design through deployment and operational support across AWS, GCP, OCI, and on-premises data centers Develop RESTful APIs and SDKs that enable other teams to easily integrate with streaming platform capabilities Work closely with platform consumers, SREs, and engineering teams across the organization to understand requirements and deliver scalable solutions Challenge the status quo by continuously improving platform performance, reliability, scalability, and developer experience Relentlessly pursue quality through comprehensive testing strategies, effective monitoring and alerting, chaos engineering practices, and resilient architecture patterns Implement observability solutions using metrics, logging, and distributed tracing to ensure platform health and performance Provide operational support and participate in on-call rotation for production streaming services Contribute to platform documentation, runbooks, and best practices for streaming application development What You'll Need Be an empathetic and a team player (Remember: One team. One fight!) Bachelor's or Master's degree in Computer Science or equivalent practical experience 8+ years of software engineering experience with proven track record of building production systems Proven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes. Strong understanding of distributed systems, stream processing concepts, event-driven architectures, and scalability patterns Proficiency in Java, Python and/or Go for building resilient, high-performance services Experience with Apache Kafka or similar distributed messaging systems (Pulsar, RabbitMQ, etc.) Hands-on experience with Apache Flink and/or Apache Spark for stream and batch processing Solid grasp of containerization and Kubernetes for deploying and orchestrating applications Proven ability to translate business requirements into technical solutions and lead projects to successful delivery Strong problem-solving skills with experience debugging complex distributed systems issues Passion for platform engineering and enabling other teams to build better products Excellent communication and collaboration skills across functions and organizational levels Ownership mindset - willingness to take initiative and drive issues to resolution Bonus Points Experience deploying and operating services across multiple cloud providers (AWS, GCP, Azure, OCI) Knowledge of data serialization formats (Avro, Protobuf, Parquet) and schema management Experience with observability platforms (Prometheus, Grafana, Datadog, Splunk) Understanding of data governance, security, and compliance requirements Contributions to open-source streaming or distributed systems projects Experience with CI/CD pipelines and GitOps practices Background in building developer platforms or internal tooling Knowledge of exactly-once processing semantics and state management in streaming systems Benefits of Working at CrowdStrike Market leader in compensation and equity awards Comprehensive physical and mental wellness programs Competitive vacation and holidays for recharge Paid parental and adoption leaves Professional development opportunities for all employees regardless of level or role Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections Vibrant office culture with world class amenities Great Place to Work Certified across the globe CrowdStrike is proud to be an equal opportunity employer. We are committed to fostering a culture of belonging where everyone is valued for who they are and empowered to succeed. We support veterans and individuals with disabilities through our affirmative action program. CrowdStrike is committed to providing equal employment opportunity for all employees and applicants for employment. The Company does not discriminate in employment opportunities or practices on the basis of race, color, creed, ethnicity, religion, sex (including pregnancy or pregnancy-related medical conditions), sexual orientation, gender identity, marital or family status, veteran status, age, national origin, ancestry, physical disability (including HIV and AIDS), mental disability, medical condition, genetic information, membership or activity in a local human rights commission, status with regard to public assistance, or any other characteristic protected by law. We base all employment decisions including recruitment, selection, training, compensation, benefits, discipline, promotions, transfers, lay-offs, return from lay-off, terminations and social/recreational programs on valid job requirements. CrowdStrike was founded in 2011 to fix a fundamental problem: The sophisticated attacks that were forcing the world's leading businesses into the headlines could not be solved with existing malware-based defenses. Founder George Kurtz realized that a brand new approach was needed - one that combines the most advanced endpoint protection with expert intelligence to pinpoint the adversaries perpetrating the attacks, not just the malware. There's much more to the story of how Falcon has redefined endpoint protection but there's only one thing to remember about CrowdStrike: We stop breaches.
26/07/2026
Full time
As a global leader in cybersecurity, CrowdStrike protects the people, processes and technologies that drive modern organizations. Since 2011, our mission hasn't changed - we're here to stop breaches, and we've redefined modern security with the world's most advanced AI-native platform. We work on large scale distributed systems, processing almost 3 trillion events per day and this traffic is growing daily. Our customers span all industries, and they count on CrowdStrike to keep their businesses running, their communities safe and their lives moving forward. We're proud to work for a mission-driven company leveraging AI to transform the way we work. CrowdStrikers drive their careers through flexibility and autonomy while also being expected to contribute to a culture of responsible AI adoption, experimentation, and innovation. We use an AI-first mindset as a force multiplier to proactively and continuously accelerate execution, build expertise, uncover insights, and solve complex problems. We're always looking to add talented CrowdStrikers to the team who have limitless passion, a relentless focus on innovation and a fanatical commitment to our customers, our community and each other. Ready to join a mission that matters? The future of cybersecurity starts with you. About the Role We are seeking a talented Senior Software Engineer to join our Real-Time Streaming Platform team. Our mission is to provide enterprise-grade streaming processing capabilities that empower teams across CrowdStrike to build real-time data pipelines, perform complex event processing, and derive actionable insights from massive data streams. You'll be building the foundational infrastructure that enables our security products to detect threats, analyze patterns, and respond to incidents in real-time. You'll be designing and developing highly scalable streaming services using technologies like Apache Kafka, Apache Flink, Apache Storm and Apache Spark in Java, Python and Go. You'll design and implement data pipelines that process trillions of events daily, ensuring low latency, high throughput, and fault tolerance. Your work will directly impact how CrowdStrike ingests, processes, and analyzes security telemetry at petabyte scale across our global customer base. Have no fear if you haven't worked with all these technologies - we value strong distributed systems fundamentals, the ability to write production-quality code, and the passion for solving complex engineering challenges. We'll support your growth in streaming technologies and expect you'll be comfortable collaborating with teams distributed across various geographies and time zones. Sounds like a challenge? It is! If that's what you're after, we'd love to hear from you. What You'll Do Design, develop, and maintain streaming microservices in Java and Go that provide real-time data processing capabilities to internal teams Build and optimize Apache Flink and Apache Spark based data pipelines running on Kubernetes infrastructure Design event-driven systems leveraging Kafka for high-throughput, low-latency message streaming Take end-to-end ownership of technical initiatives, from design through deployment and operational support across AWS, GCP, OCI, and on-premises data centers Develop RESTful APIs and SDKs that enable other teams to easily integrate with streaming platform capabilities Work closely with platform consumers, SREs, and engineering teams across the organization to understand requirements and deliver scalable solutions Challenge the status quo by continuously improving platform performance, reliability, scalability, and developer experience Relentlessly pursue quality through comprehensive testing strategies, effective monitoring and alerting, chaos engineering practices, and resilient architecture patterns Implement observability solutions using metrics, logging, and distributed tracing to ensure platform health and performance Provide operational support and participate in on-call rotation for production streaming services Contribute to platform documentation, runbooks, and best practices for streaming application development What You'll Need Be an empathetic and a team player (Remember: One team. One fight!) Bachelor's or Master's degree in Computer Science or equivalent practical experience 8+ years of software engineering experience with proven track record of building production systems Proven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes. Strong understanding of distributed systems, stream processing concepts, event-driven architectures, and scalability patterns Proficiency in Java, Python and/or Go for building resilient, high-performance services Experience with Apache Kafka or similar distributed messaging systems (Pulsar, RabbitMQ, etc.) Hands-on experience with Apache Flink and/or Apache Spark for stream and batch processing Solid grasp of containerization and Kubernetes for deploying and orchestrating applications Proven ability to translate business requirements into technical solutions and lead projects to successful delivery Strong problem-solving skills with experience debugging complex distributed systems issues Passion for platform engineering and enabling other teams to build better products Excellent communication and collaboration skills across functions and organizational levels Ownership mindset - willingness to take initiative and drive issues to resolution Bonus Points Experience deploying and operating services across multiple cloud providers (AWS, GCP, Azure, OCI) Knowledge of data serialization formats (Avro, Protobuf, Parquet) and schema management Experience with observability platforms (Prometheus, Grafana, Datadog, Splunk) Understanding of data governance, security, and compliance requirements Contributions to open-source streaming or distributed systems projects Experience with CI/CD pipelines and GitOps practices Background in building developer platforms or internal tooling Knowledge of exactly-once processing semantics and state management in streaming systems Benefits of Working at CrowdStrike Market leader in compensation and equity awards Comprehensive physical and mental wellness programs Competitive vacation and holidays for recharge Paid parental and adoption leaves Professional development opportunities for all employees regardless of level or role Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections Vibrant office culture with world class amenities Great Place to Work Certified across the globe CrowdStrike is proud to be an equal opportunity employer. We are committed to fostering a culture of belonging where everyone is valued for who they are and empowered to succeed. We support veterans and individuals with disabilities through our affirmative action program. CrowdStrike is committed to providing equal employment opportunity for all employees and applicants for employment. The Company does not discriminate in employment opportunities or practices on the basis of race, color, creed, ethnicity, religion, sex (including pregnancy or pregnancy-related medical conditions), sexual orientation, gender identity, marital or family status, veteran status, age, national origin, ancestry, physical disability (including HIV and AIDS), mental disability, medical condition, genetic information, membership or activity in a local human rights commission, status with regard to public assistance, or any other characteristic protected by law. We base all employment decisions including recruitment, selection, training, compensation, benefits, discipline, promotions, transfers, lay-offs, return from lay-off, terminations and social/recreational programs on valid job requirements. CrowdStrike was founded in 2011 to fix a fundamental problem: The sophisticated attacks that were forcing the world's leading businesses into the headlines could not be solved with existing malware-based defenses. Founder George Kurtz realized that a brand new approach was needed - one that combines the most advanced endpoint protection with expert intelligence to pinpoint the adversaries perpetrating the attacks, not just the malware. There's much more to the story of how Falcon has redefined endpoint protection but there's only one thing to remember about CrowdStrike: We stop breaches.
The position works closely with IAM architects, business stakeholders and technology partners. In this role, you will: Build and maintain cloud-native platform services supporting the IDAM 2.0 programme. Design, deploy and manage Kubernetes clusters and supporting platform components. Develop Infrastructure as Code (IaC) for repeatable, automated deployments. Implement and maintain CI/CD pipelines and GitOps deployment workflows. Manage cloud networking, connectivity and platform security. Implement platform observability including logging, monitoring, metrics and distributed tracing. Automate platform provisioning, configuration management and operational tasks. Support deployment and operation of identity platform components and supporting services. Implement secure secrets management and certificate lifecycle automation. Configure service mesh technologies and secure service-to-service communication. Manage ingress, API gateways and network routing components. Implement platform resilience, backup, disaster recovery and high availability capabilities. Support workload identity and platform authentication mechanisms. Ensure platform compliance with enterprise security, governance and regulatory requirements. Optimise platform performance, scalability and cost efficiency. Develop operational tooling, scripts and self-service capabilities. Support incident management, troubleshooting and root cause analysis. Collaborate with Security, Architecture and Engineering teams to deliver secure platform services. Contribute to platform standards, operational documentation and runbooks. Support environment management across development, test, pilot and production environments. Drive continuous improvement through automation and platform engineering best practices. To be successful in this role, you should meet the following requirements: Key Skills & Experience Essential Skills Cloud Platform Engineering (GCP) Kubernetes Administration Infrastructure as Code (Terraform or equivalent) CI/CD Pipelines GitOps Linux Administration Networking and Load Balancing Service Mesh Technologies (Istio/Envoy) Container Platforms Secrets Management PKI and Certificate Management Workload Identity Observability (Logging, Monitoring and Tracing) Scripting and Automation DevSecOps Site Reliability Engineering (SRE) Performance and Capacity Management Operational Support Global Deployment Strategies Positive can-do attitude adept at solutionising in large, complex organisations GCS is acting as an Employment Business in relation to this vacancy.
25/07/2026
Contractor
The position works closely with IAM architects, business stakeholders and technology partners. In this role, you will: Build and maintain cloud-native platform services supporting the IDAM 2.0 programme. Design, deploy and manage Kubernetes clusters and supporting platform components. Develop Infrastructure as Code (IaC) for repeatable, automated deployments. Implement and maintain CI/CD pipelines and GitOps deployment workflows. Manage cloud networking, connectivity and platform security. Implement platform observability including logging, monitoring, metrics and distributed tracing. Automate platform provisioning, configuration management and operational tasks. Support deployment and operation of identity platform components and supporting services. Implement secure secrets management and certificate lifecycle automation. Configure service mesh technologies and secure service-to-service communication. Manage ingress, API gateways and network routing components. Implement platform resilience, backup, disaster recovery and high availability capabilities. Support workload identity and platform authentication mechanisms. Ensure platform compliance with enterprise security, governance and regulatory requirements. Optimise platform performance, scalability and cost efficiency. Develop operational tooling, scripts and self-service capabilities. Support incident management, troubleshooting and root cause analysis. Collaborate with Security, Architecture and Engineering teams to deliver secure platform services. Contribute to platform standards, operational documentation and runbooks. Support environment management across development, test, pilot and production environments. Drive continuous improvement through automation and platform engineering best practices. To be successful in this role, you should meet the following requirements: Key Skills & Experience Essential Skills Cloud Platform Engineering (GCP) Kubernetes Administration Infrastructure as Code (Terraform or equivalent) CI/CD Pipelines GitOps Linux Administration Networking and Load Balancing Service Mesh Technologies (Istio/Envoy) Container Platforms Secrets Management PKI and Certificate Management Workload Identity Observability (Logging, Monitoring and Tracing) Scripting and Automation DevSecOps Site Reliability Engineering (SRE) Performance and Capacity Management Operational Support Global Deployment Strategies Positive can-do attitude adept at solutionising in large, complex organisations GCS is acting as an Employment Business in relation to this vacancy.
We're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people centric and here to power good. Every day, we future proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our company and customers from what's now to what's next. We make it happen through the power of acceleration. Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us. Job Description Mandatory Skills: Plat Engineering for on premises & cloud, IaC, CI/CD pipeline & tools, security, network, landing zones, observability, self servicing, DB automation ROLE PURPOSE Lead the design, build, and operational excellence of the organisation's cloud platform across private and public cloud environments. Own the end to end IaC based provisioning, configuration management, patch orchestration, and cloud governance framework. Drive the transformation from manual infrastructure operations to a fully automated, self service, policy governed cloud platform that delivers speed, security, cost efficiency, and reliability. KEY RESPONSIBILITIES Architect and implement IaC based cloud provisioning using Terraform, CloudFormation, or Pulumi across multi cloud and private cloud environments Design and automate end to end infrastructure lifecycle management: provisioning, patching, updates, scaling, and decommissioning Build and maintain cloud governance guardrails using policy as code (OPA, Sentinel, Azure Policy) to enforce security, cost, and compliance standards Implement FinOps practices including automated rightsizing, reserved instance management, waste detection, and cost allocation tagging Lead CI/CD pipeline engineering for infrastructure deployments, integrating automated testing, security scanning, and approval gates Design self service infrastructure portals enabling development teams to provision compliant environments without manual intervention Mentor and coach 2 Cloud Platform Engineers, establish coding standards, conduct architecture reviews, and drive knowledge sharing Collaborate with SRE and Security teams to embed reliability, observability, and compliance into the platform from the ground up Produce architectural decision records (ADRs), runbooks, and platform documentation for operational excellence TECHNICAL SKILLS & EXPERTISE Expert level IaC proficiency: Terraform (strongly preferred), CloudFormation, Pulumi, or Bicep Deep experience with at least two major cloud platforms: AWS, Azure, GCP, or VMware private cloud Strong experience with Kubernetes (EKS/AKS/GKE), container orchestration, and platform services Configuration management expertise: Ansible, Chef, Puppet, or SaltStack CI/CD tooling: Jenkins, GitLab CI, GitHub Actions, Azure DevOps for infrastructure pipelines Policy as code frameworks: OPA/Rego, HashiCorp Sentinel, Azure Policy, AWS Config Rules Monitoring and observability: Prometheus, Grafana, Datadog, CloudWatch, or Dynatrace Networking fundamentals: VPC/VNet design, load balancers, DNS, CDN, and hybrid connectivity SOFT SKILLS & COMPETENCIES Strong leadership and mentoring ability - coaches and develops junior engineers Excellent stakeholder management and communication skills across all levels Strategic thinker who balances technical depth with business outcomes Proven ability to drive change, influence without authority, and build consensus Strong analytical and problem solving mindset with attention to detail Ability to manage competing priorities across multiple workstreams simultaneously QUALIFICATIONS & EXPERIENCE 7+ years in cloud/infrastructure engineering with 3+ years in a senior or lead capacity Proven track record of delivering large scale IaC transformation and cloud migration programmes Relevant certifications: AWS Solutions Architect Professional, Azure Solutions Architect Expert, CKA/CKAD, or Terraform Associate/Professional Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience) DESIRABLE / NICE TO HAVE Experience with FinOps frameworks and cloud cost optimisation at scale Background in platform engineering product management and internal developer platforms Experience with GitOps workflows (ArgoCD, Flux) and service mesh (Istio, Linkerd) About Us We're a global, team of innovators. Together, we harness engineering excellence and passion to co create meaningful solutions to complex challenges. We turn organisations into data driven leaders that can make a positive impact on their industries and society. If you believe that innovation can bring a better tomorrow closer to today, this is the place for you. Fostering innovation through diverse perspectives Hitachi is a global company operating across a wide range of industries and regions. One of the things that sets Hitachi apart is the diversity of our business and people, which drives our innovation and growth. We are committed to building an inclusive culture based on mutual respect and merit based systems. We believe that when people feel valued, heard, and safe to express themselves, they do their best work. How we look after you We help take care of your today and tomorrow with industry leading benefits, support, and services that look after your holistic health and wellbeing. We're also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We're always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you'll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with. We're proud to say we're an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
24/07/2026
Full time
We're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people centric and here to power good. Every day, we future proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our company and customers from what's now to what's next. We make it happen through the power of acceleration. Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us. Job Description Mandatory Skills: Plat Engineering for on premises & cloud, IaC, CI/CD pipeline & tools, security, network, landing zones, observability, self servicing, DB automation ROLE PURPOSE Lead the design, build, and operational excellence of the organisation's cloud platform across private and public cloud environments. Own the end to end IaC based provisioning, configuration management, patch orchestration, and cloud governance framework. Drive the transformation from manual infrastructure operations to a fully automated, self service, policy governed cloud platform that delivers speed, security, cost efficiency, and reliability. KEY RESPONSIBILITIES Architect and implement IaC based cloud provisioning using Terraform, CloudFormation, or Pulumi across multi cloud and private cloud environments Design and automate end to end infrastructure lifecycle management: provisioning, patching, updates, scaling, and decommissioning Build and maintain cloud governance guardrails using policy as code (OPA, Sentinel, Azure Policy) to enforce security, cost, and compliance standards Implement FinOps practices including automated rightsizing, reserved instance management, waste detection, and cost allocation tagging Lead CI/CD pipeline engineering for infrastructure deployments, integrating automated testing, security scanning, and approval gates Design self service infrastructure portals enabling development teams to provision compliant environments without manual intervention Mentor and coach 2 Cloud Platform Engineers, establish coding standards, conduct architecture reviews, and drive knowledge sharing Collaborate with SRE and Security teams to embed reliability, observability, and compliance into the platform from the ground up Produce architectural decision records (ADRs), runbooks, and platform documentation for operational excellence TECHNICAL SKILLS & EXPERTISE Expert level IaC proficiency: Terraform (strongly preferred), CloudFormation, Pulumi, or Bicep Deep experience with at least two major cloud platforms: AWS, Azure, GCP, or VMware private cloud Strong experience with Kubernetes (EKS/AKS/GKE), container orchestration, and platform services Configuration management expertise: Ansible, Chef, Puppet, or SaltStack CI/CD tooling: Jenkins, GitLab CI, GitHub Actions, Azure DevOps for infrastructure pipelines Policy as code frameworks: OPA/Rego, HashiCorp Sentinel, Azure Policy, AWS Config Rules Monitoring and observability: Prometheus, Grafana, Datadog, CloudWatch, or Dynatrace Networking fundamentals: VPC/VNet design, load balancers, DNS, CDN, and hybrid connectivity SOFT SKILLS & COMPETENCIES Strong leadership and mentoring ability - coaches and develops junior engineers Excellent stakeholder management and communication skills across all levels Strategic thinker who balances technical depth with business outcomes Proven ability to drive change, influence without authority, and build consensus Strong analytical and problem solving mindset with attention to detail Ability to manage competing priorities across multiple workstreams simultaneously QUALIFICATIONS & EXPERIENCE 7+ years in cloud/infrastructure engineering with 3+ years in a senior or lead capacity Proven track record of delivering large scale IaC transformation and cloud migration programmes Relevant certifications: AWS Solutions Architect Professional, Azure Solutions Architect Expert, CKA/CKAD, or Terraform Associate/Professional Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience) DESIRABLE / NICE TO HAVE Experience with FinOps frameworks and cloud cost optimisation at scale Background in platform engineering product management and internal developer platforms Experience with GitOps workflows (ArgoCD, Flux) and service mesh (Istio, Linkerd) About Us We're a global, team of innovators. Together, we harness engineering excellence and passion to co create meaningful solutions to complex challenges. We turn organisations into data driven leaders that can make a positive impact on their industries and society. If you believe that innovation can bring a better tomorrow closer to today, this is the place for you. Fostering innovation through diverse perspectives Hitachi is a global company operating across a wide range of industries and regions. One of the things that sets Hitachi apart is the diversity of our business and people, which drives our innovation and growth. We are committed to building an inclusive culture based on mutual respect and merit based systems. We believe that when people feel valued, heard, and safe to express themselves, they do their best work. How we look after you We help take care of your today and tomorrow with industry leading benefits, support, and services that look after your holistic health and wellbeing. We're also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We're always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you'll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with. We're proud to say we're an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
Function Cloud & Data Engineering About the Company We're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people centric and here to power good. Every day, we future prove urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our company and customers from what's now to what's next. We make it happen through the power of acceleration. Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us. Job Description Mandatory Skills Plat Engineering for on premises & cloud, IaC, CI/CD pipeline & tools, security, network, landing zones, observability, self servicing, DB automation Role Purpose Lead the design, build, and operational excellence of the organisation's cloud platform across private and public cloud environments. Own the end to end IaC based provisioning, configuration management, patch orchestration, and cloud governance framework. Drive the transformation from manual infrastructure operations to a fully automated, self service, policy governed cloud platform that delivers speed, security, cost efficiency, and reliability. Key Responsibilities Architect and implement IaC based cloud provisioning using Terraform, CloudFormation, or Pulumi across multi cloud and private cloud environments Design and automate end to end infrastructure lifecycle management: provisioning, patching, updates, scaling, and decommissioning Build and maintain cloud governance guardrails using policy as code (OPA, Sentinel, Azure Policy) to enforce security, cost, and compliance standards Implement FinOps practices including automated rightsizing, reserved instance management, waste detection, and cost allocation tagging Lead CI/CD pipeline engineering for infrastructure deployments, integrating automated testing, security scanning, and approval gates Design self service infrastructure portals enabling development teams to provision compliant environments without manual intervention Mentor and coach 2 Cloud Platform Engineers, establish coding standards, conduct architecture reviews, and drive knowledge sharing Collaborate with SRE and Security teams to embed reliability, observability, and compliance into the platform from the ground up Produce architectural decision records (ADRs), runbooks, and platform documentation for operational excellence Technical Skills & Expertise Expert level IaC proficiency: Terraform (strongly preferred), CloudFormation, Pulumi, or Bicep Deep experience with at least two major cloud platforms: AWS, Azure, GCP, or VMware private cloud Strong experience with Kubernetes (EKS/AKS/GKE), container orchestration, and platform services Configuration management expertise: Ansible, Chef, Puppet, or SaltStack CI/CD tooling: Jenkins, GitLab CI, GitHub Actions, Azure DevOps for infrastructure pipelines Policy as code frameworks: OPA/Rego, HashiCorp Sentinel, Azure Policy, AWS Config Rules Monitoring and observability: Prometheus, Grafana, Datadog, CloudWatch, or Dynatrace Networking fundamentals: VPC/VNet design, load balancers, DNS, CDN, and hybrid connectivity Soft Skills & Competencies Strong leadership and mentoring ability - coaches and develops junior engineers Excellent stakeholder management and communication skills across all levels Strategic thinker who balances technical depth with business outcomes Proven ability to drive change, influence without authority, and build consensus Strong analytical and problem solving mindset with attention to detail Ability to manage competing priorities across multiple workstreams simultaneously Qualifications & Experience 7+ years in cloud/infrastructure engineering with 3+ years in a senior or lead capacity Proven track record of delivering large scale IaC transformation and cloud migration programmes Relevant certifications: AWS Solutions Architect Professional, Azure Solutions Architect Expert, CKA/CKAD, or Terraform Associate/Professional Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience) Desirable / Nice to Have Experience with FinOps frameworks and cloud cost optimisation at scale Background in platform engineering product management and internal developer platforms Experience with GitOps workflows (ArgoCD, Flux) and service mesh (Istio, Linkerd) Benefits & Well being We help take care of your today and tomorrow with industry leading benefits, support, and services that look after your holistic health and wellbeing. We're also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We're always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you'll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with. Equal Opportunity Statement We're proud to say we're an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
24/07/2026
Full time
Function Cloud & Data Engineering About the Company We're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people centric and here to power good. Every day, we future prove urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our company and customers from what's now to what's next. We make it happen through the power of acceleration. Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us. Job Description Mandatory Skills Plat Engineering for on premises & cloud, IaC, CI/CD pipeline & tools, security, network, landing zones, observability, self servicing, DB automation Role Purpose Lead the design, build, and operational excellence of the organisation's cloud platform across private and public cloud environments. Own the end to end IaC based provisioning, configuration management, patch orchestration, and cloud governance framework. Drive the transformation from manual infrastructure operations to a fully automated, self service, policy governed cloud platform that delivers speed, security, cost efficiency, and reliability. Key Responsibilities Architect and implement IaC based cloud provisioning using Terraform, CloudFormation, or Pulumi across multi cloud and private cloud environments Design and automate end to end infrastructure lifecycle management: provisioning, patching, updates, scaling, and decommissioning Build and maintain cloud governance guardrails using policy as code (OPA, Sentinel, Azure Policy) to enforce security, cost, and compliance standards Implement FinOps practices including automated rightsizing, reserved instance management, waste detection, and cost allocation tagging Lead CI/CD pipeline engineering for infrastructure deployments, integrating automated testing, security scanning, and approval gates Design self service infrastructure portals enabling development teams to provision compliant environments without manual intervention Mentor and coach 2 Cloud Platform Engineers, establish coding standards, conduct architecture reviews, and drive knowledge sharing Collaborate with SRE and Security teams to embed reliability, observability, and compliance into the platform from the ground up Produce architectural decision records (ADRs), runbooks, and platform documentation for operational excellence Technical Skills & Expertise Expert level IaC proficiency: Terraform (strongly preferred), CloudFormation, Pulumi, or Bicep Deep experience with at least two major cloud platforms: AWS, Azure, GCP, or VMware private cloud Strong experience with Kubernetes (EKS/AKS/GKE), container orchestration, and platform services Configuration management expertise: Ansible, Chef, Puppet, or SaltStack CI/CD tooling: Jenkins, GitLab CI, GitHub Actions, Azure DevOps for infrastructure pipelines Policy as code frameworks: OPA/Rego, HashiCorp Sentinel, Azure Policy, AWS Config Rules Monitoring and observability: Prometheus, Grafana, Datadog, CloudWatch, or Dynatrace Networking fundamentals: VPC/VNet design, load balancers, DNS, CDN, and hybrid connectivity Soft Skills & Competencies Strong leadership and mentoring ability - coaches and develops junior engineers Excellent stakeholder management and communication skills across all levels Strategic thinker who balances technical depth with business outcomes Proven ability to drive change, influence without authority, and build consensus Strong analytical and problem solving mindset with attention to detail Ability to manage competing priorities across multiple workstreams simultaneously Qualifications & Experience 7+ years in cloud/infrastructure engineering with 3+ years in a senior or lead capacity Proven track record of delivering large scale IaC transformation and cloud migration programmes Relevant certifications: AWS Solutions Architect Professional, Azure Solutions Architect Expert, CKA/CKAD, or Terraform Associate/Professional Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience) Desirable / Nice to Have Experience with FinOps frameworks and cloud cost optimisation at scale Background in platform engineering product management and internal developer platforms Experience with GitOps workflows (ArgoCD, Flux) and service mesh (Istio, Linkerd) Benefits & Well being We help take care of your today and tomorrow with industry leading benefits, support, and services that look after your holistic health and wellbeing. We're also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We're always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you'll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with. Equal Opportunity Statement We're proud to say we're an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
At CoMind, we are developing a non-invasive neuromonitoring technology that will result in a new era of clinical brain monitoring. In joining us, you will be helping to create cutting-edge technologies that will improve how we diagnose and treat brain disorders, ultimately improving and saving the lives of patients across the world. The Role: The Software Engineering team at CoMind spans the full stack of a regulated medical device company from embedded and real-time systems through application software, ML inference, GUI, DevOps, and test engineering. It is a team that builds across every layer, working to medical device software standards and shipping software that has to perform reliably in clinical environments. As a Senior/Staff Software Test Engineer, you will define the quality and verification strategy across CoMind's entire software stack from embedded real-time systems to cloud-side data pipelines. You will move beyond task-level execution to architect the infrastructure and 'compliance-as-code' processes that provide the engineering team with high-velocity, high-confidence validation. You will serve as a key technical authority, ensuring our software not only meets stringent IEC 62304 and ISO 14971 standards but does so through sustainable, automated design. You will function as the technical authority for software quality. Your goal is to move CoMind beyond traditional V&V, architecting a self-validating system that integrates safety, regulatory compliance, and high-velocity development. You will be the bridge between Systems Engineering, Regulatory Affairs, and Software Engineering, ensuring that quality is an inherent attribute of our technology rather than a post-development checklist. At CoMind, all team members work at least 4 days per week in the office, plus a flexible work-from-home day. This role is based in our London (Kings Cross) office. Responsibilities: Architect a scalable, high-fidelity testing infrastructure that enables continuous validation. Champion 'design for testability' across the entire engineering stack, influencing system architecture to ensure critical safety and performance requirements are verifiable by design. Define the end-to-end V&V strategy for the product lifecycle. Quantitatively map software verification evidence directly to system-level safety risks (ISO 14971) to create a 'compliance-as-code' environment that drastically reduces time-to-submission for FDA and other regulatory bodies. Drive a 'quality-first' engineering culture by mentorship. Lead training and initiatives on Test-Driven Development (TDD), CI/CD best practices, and automated quality gates, ensuring that the engineering team's velocity is sustained by high-confidence automated processes. Lead the development of AI-augmented test strategies. Architect predictive anomaly detection systems and AI-driven test generation engines that evolve alongside our software, turning the test suite into an intelligent system that learns from past regressions. Collaborate with engineering leadership to drive continuous improvement in software development and test processes, ensuring that quality practices scale with the organisation. AI is fundamental to our culture - it's not just a tool, but a core part of how we work, collaborate, and innovate. We expect all team members to embrace AI in their daily work and continuously find new ways to use it effectively. Skills & Experience: Strategic Execution: 7+ years in software test engineering, with a proven track record of architecting testing strategies for complex, regulated systems (embedded, real-time, or SaMD). Regulatory & Risk Authority: Deep working knowledge of IEC 62304 and ISO 14971. Proven ability to design 'compliance-as-code' workflows that integrate risk controls directly into the CI/CD pipeline and automate regulatory evidence generation. Technical Proficiency: Advanced Python and C/C++; extensive experience building scalable test infrastructure for both embedded hardware and cloud-side data pipelines. Cultural Leadership: Demonstrated success in mentoring software engineers and driving a 'quality-first' culture through TDD, static analysis, and automated quality gates. Cybersecurity & Documentation: Strong experience in cybersecurity requirements (e.g., IEC 81001-5-1, FDA cybersecurity guidance) and rigorous documentation management for regulatory submission. Nice to Have Experience testing real-time Linux or RTOS-based products. Background in MLOps, including validation of model drift, dataset governance, and automated performance evaluation. Expertise in managing complex build environments (e.g., CMake, Bazel) or cloud-native infrastructure-as-code (Terraform, AWS/GCP). Direct experience navigating FDA Q-Sub or MHRA pre-submission processes. Experience with hardware-in-the-loop or software-in-the-loop test environments. Benefits: Company equity plan so all employees share in the success of the company Salary-sacrifice pension scheme Private medical, dental and vision insurance (medical history disregarded) Group life assurance at 4x annual income Comprehensive mental health support, including unlimited access to 1:1 sessions with trained professionals Unlimited holiday allowance (+ bank holidays) and one week of remote working per quarter Lunch voucher (£10) every day for JustEat and free dinner on those days where you need to work later Twice weekly deliveries of fresh fruit and an extensive selection of snacks and drinks YuLife subscription, allowing you to turn your daily steps and meditation into discounts at a range of stores
22/07/2026
Full time
At CoMind, we are developing a non-invasive neuromonitoring technology that will result in a new era of clinical brain monitoring. In joining us, you will be helping to create cutting-edge technologies that will improve how we diagnose and treat brain disorders, ultimately improving and saving the lives of patients across the world. The Role: The Software Engineering team at CoMind spans the full stack of a regulated medical device company from embedded and real-time systems through application software, ML inference, GUI, DevOps, and test engineering. It is a team that builds across every layer, working to medical device software standards and shipping software that has to perform reliably in clinical environments. As a Senior/Staff Software Test Engineer, you will define the quality and verification strategy across CoMind's entire software stack from embedded real-time systems to cloud-side data pipelines. You will move beyond task-level execution to architect the infrastructure and 'compliance-as-code' processes that provide the engineering team with high-velocity, high-confidence validation. You will serve as a key technical authority, ensuring our software not only meets stringent IEC 62304 and ISO 14971 standards but does so through sustainable, automated design. You will function as the technical authority for software quality. Your goal is to move CoMind beyond traditional V&V, architecting a self-validating system that integrates safety, regulatory compliance, and high-velocity development. You will be the bridge between Systems Engineering, Regulatory Affairs, and Software Engineering, ensuring that quality is an inherent attribute of our technology rather than a post-development checklist. At CoMind, all team members work at least 4 days per week in the office, plus a flexible work-from-home day. This role is based in our London (Kings Cross) office. Responsibilities: Architect a scalable, high-fidelity testing infrastructure that enables continuous validation. Champion 'design for testability' across the entire engineering stack, influencing system architecture to ensure critical safety and performance requirements are verifiable by design. Define the end-to-end V&V strategy for the product lifecycle. Quantitatively map software verification evidence directly to system-level safety risks (ISO 14971) to create a 'compliance-as-code' environment that drastically reduces time-to-submission for FDA and other regulatory bodies. Drive a 'quality-first' engineering culture by mentorship. Lead training and initiatives on Test-Driven Development (TDD), CI/CD best practices, and automated quality gates, ensuring that the engineering team's velocity is sustained by high-confidence automated processes. Lead the development of AI-augmented test strategies. Architect predictive anomaly detection systems and AI-driven test generation engines that evolve alongside our software, turning the test suite into an intelligent system that learns from past regressions. Collaborate with engineering leadership to drive continuous improvement in software development and test processes, ensuring that quality practices scale with the organisation. AI is fundamental to our culture - it's not just a tool, but a core part of how we work, collaborate, and innovate. We expect all team members to embrace AI in their daily work and continuously find new ways to use it effectively. Skills & Experience: Strategic Execution: 7+ years in software test engineering, with a proven track record of architecting testing strategies for complex, regulated systems (embedded, real-time, or SaMD). Regulatory & Risk Authority: Deep working knowledge of IEC 62304 and ISO 14971. Proven ability to design 'compliance-as-code' workflows that integrate risk controls directly into the CI/CD pipeline and automate regulatory evidence generation. Technical Proficiency: Advanced Python and C/C++; extensive experience building scalable test infrastructure for both embedded hardware and cloud-side data pipelines. Cultural Leadership: Demonstrated success in mentoring software engineers and driving a 'quality-first' culture through TDD, static analysis, and automated quality gates. Cybersecurity & Documentation: Strong experience in cybersecurity requirements (e.g., IEC 81001-5-1, FDA cybersecurity guidance) and rigorous documentation management for regulatory submission. Nice to Have Experience testing real-time Linux or RTOS-based products. Background in MLOps, including validation of model drift, dataset governance, and automated performance evaluation. Expertise in managing complex build environments (e.g., CMake, Bazel) or cloud-native infrastructure-as-code (Terraform, AWS/GCP). Direct experience navigating FDA Q-Sub or MHRA pre-submission processes. Experience with hardware-in-the-loop or software-in-the-loop test environments. Benefits: Company equity plan so all employees share in the success of the company Salary-sacrifice pension scheme Private medical, dental and vision insurance (medical history disregarded) Group life assurance at 4x annual income Comprehensive mental health support, including unlimited access to 1:1 sessions with trained professionals Unlimited holiday allowance (+ bank holidays) and one week of remote working per quarter Lunch voucher (£10) every day for JustEat and free dinner on those days where you need to work later Twice weekly deliveries of fresh fruit and an extensive selection of snacks and drinks YuLife subscription, allowing you to turn your daily steps and meditation into discounts at a range of stores
At CoMind, we are developing a non-invasive neuromonitoring technology that will result in a new era of clinical brain monitoring. In joining us, you will be helping to create cutting-edge technologies that will improve how we diagnose and treat brain disorders, ultimately improving and saving the lives of patients across the world. The Role: The Software Engineering team at CoMind spans the full stack of a regulated medical device company from embedded and real-time systems through application software, ML inference, GUI, DevOps, and test engineering. It is a team that builds across every layer, working to medical device software standards and shipping software that has to perform reliably in clinical environments. As a Senior/Staff Software Test Engineer, you will define the quality and verification strategy across CoMind's entire software stack from embedded real-time systems to cloud-side data pipelines. You will move beyond task-level execution to architect the infrastructure and 'compliance-as-code' processes that provide the engineering team with high-velocity, high-confidence validation. You will serve as a key technical authority, ensuring our software not only meets stringent IEC 62304 and ISO 14971 standards but does so through sustainable, automated design. You will function as the technical authority for software quality. Your goal is to move CoMind beyond traditional V&V, architecting a self-validating system that integrates safety, regulatory compliance, and high-velocity development. You will be the bridge between Systems Engineering, Regulatory Affairs, and Software Engineering, ensuring that quality is an inherent attribute of our technology rather than a post-development checklist. At CoMind, all team members work at least 4 days per week in the office, plus a flexible work-from-home day. This role is based in our London (Kings Cross) office. Responsibilities: Architect a scalable, high-fidelity testing infrastructure that enables continuous validation. Champion 'design for testability' across the entire engineering stack, influencing system architecture to ensure critical safety and performance requirements are verifiable by design. Define the end-to-end V&V strategy for the product lifecycle. Quantitatively map software verification evidence directly to system-level safety risks (ISO 14971) to create a 'compliance-as-code' environment that drastically reduces time-to-submission for FDA and other regulatory bodies. Drive a 'quality-first' engineering culture by mentorship. Lead training and initiatives on Test-Driven Development (TDD), CI/CD best practices, and automated quality gates, ensuring that the engineering team's velocity is sustained by high-confidence automated processes. Lead the development of AI-augmented test strategies. Architect predictive anomaly detection systems and AI-driven test generation engines that evolve alongside our software, turning the test suite into an intelligent system that learns from past regressions. Collaborate with engineering leadership to drive continuous improvement in software development and test processes, ensuring that quality practices scale with the organisation. AI is fundamental to our culture - it's not just a tool, but a core part of how we work, collaborate, and innovate. We expect all team members to embrace AI in their daily work and continuously find new ways to use it effectively. Skills & Experience: Strategic Execution: 7+ years in software test engineering, with a proven track record of architecting testing strategies for complex, regulated systems (embedded, real-time, or SaMD). Regulatory & Risk Authority: Deep working knowledge of IEC 62304 and ISO 14971. Proven ability to design 'compliance-as-code' workflows that integrate risk controls directly into the CI/CD pipeline and automate regulatory evidence generation. Technical Proficiency: Advanced Python and C/C++; extensive experience building scalable test infrastructure for both embedded hardware and cloud-side data pipelines. Cultural Leadership: Demonstrated success in mentoring software engineers and driving a 'quality-first' culture through TDD, static analysis, and automated quality gates. Cybersecurity & Documentation: Strong experience in cybersecurity requirements (e.g., IEC 81001-5-1, FDA cybersecurity guidance) and rigorous documentation management for regulatory submission. Nice to Have Experience testing real-time Linux or RTOS-based products. Background in MLOps, including validation of model drift, dataset governance, and automated performance evaluation. Expertise in managing complex build environments (e.g., CMake, Bazel) or cloud-native infrastructure-as-code (Terraform, AWS/GCP). Direct experience navigating FDA Q-Sub or MHRA pre-submission processes. Experience with hardware-in-the-loop or software-in-the-loop test environments. Benefits: Company equity plan so all employees share in the success of the company Salary-sacrifice pension scheme Private medical, dental and vision insurance (medical history disregarded) Group life assurance at 4x annual income Comprehensive mental health support, including unlimited access to 1:1 sessions with trained professionals Unlimited holiday allowance (+ bank holidays) and one week of remote working per quarter Lunch voucher (£10) every day for JustEat and free dinner on those days where you need to work later Twice weekly deliveries of fresh fruit and an extensive selection of snacks and drinks YuLife subscription, allowing you to turn your daily steps and meditation into discounts at a range of stores
22/07/2026
Full time
At CoMind, we are developing a non-invasive neuromonitoring technology that will result in a new era of clinical brain monitoring. In joining us, you will be helping to create cutting-edge technologies that will improve how we diagnose and treat brain disorders, ultimately improving and saving the lives of patients across the world. The Role: The Software Engineering team at CoMind spans the full stack of a regulated medical device company from embedded and real-time systems through application software, ML inference, GUI, DevOps, and test engineering. It is a team that builds across every layer, working to medical device software standards and shipping software that has to perform reliably in clinical environments. As a Senior/Staff Software Test Engineer, you will define the quality and verification strategy across CoMind's entire software stack from embedded real-time systems to cloud-side data pipelines. You will move beyond task-level execution to architect the infrastructure and 'compliance-as-code' processes that provide the engineering team with high-velocity, high-confidence validation. You will serve as a key technical authority, ensuring our software not only meets stringent IEC 62304 and ISO 14971 standards but does so through sustainable, automated design. You will function as the technical authority for software quality. Your goal is to move CoMind beyond traditional V&V, architecting a self-validating system that integrates safety, regulatory compliance, and high-velocity development. You will be the bridge between Systems Engineering, Regulatory Affairs, and Software Engineering, ensuring that quality is an inherent attribute of our technology rather than a post-development checklist. At CoMind, all team members work at least 4 days per week in the office, plus a flexible work-from-home day. This role is based in our London (Kings Cross) office. Responsibilities: Architect a scalable, high-fidelity testing infrastructure that enables continuous validation. Champion 'design for testability' across the entire engineering stack, influencing system architecture to ensure critical safety and performance requirements are verifiable by design. Define the end-to-end V&V strategy for the product lifecycle. Quantitatively map software verification evidence directly to system-level safety risks (ISO 14971) to create a 'compliance-as-code' environment that drastically reduces time-to-submission for FDA and other regulatory bodies. Drive a 'quality-first' engineering culture by mentorship. Lead training and initiatives on Test-Driven Development (TDD), CI/CD best practices, and automated quality gates, ensuring that the engineering team's velocity is sustained by high-confidence automated processes. Lead the development of AI-augmented test strategies. Architect predictive anomaly detection systems and AI-driven test generation engines that evolve alongside our software, turning the test suite into an intelligent system that learns from past regressions. Collaborate with engineering leadership to drive continuous improvement in software development and test processes, ensuring that quality practices scale with the organisation. AI is fundamental to our culture - it's not just a tool, but a core part of how we work, collaborate, and innovate. We expect all team members to embrace AI in their daily work and continuously find new ways to use it effectively. Skills & Experience: Strategic Execution: 7+ years in software test engineering, with a proven track record of architecting testing strategies for complex, regulated systems (embedded, real-time, or SaMD). Regulatory & Risk Authority: Deep working knowledge of IEC 62304 and ISO 14971. Proven ability to design 'compliance-as-code' workflows that integrate risk controls directly into the CI/CD pipeline and automate regulatory evidence generation. Technical Proficiency: Advanced Python and C/C++; extensive experience building scalable test infrastructure for both embedded hardware and cloud-side data pipelines. Cultural Leadership: Demonstrated success in mentoring software engineers and driving a 'quality-first' culture through TDD, static analysis, and automated quality gates. Cybersecurity & Documentation: Strong experience in cybersecurity requirements (e.g., IEC 81001-5-1, FDA cybersecurity guidance) and rigorous documentation management for regulatory submission. Nice to Have Experience testing real-time Linux or RTOS-based products. Background in MLOps, including validation of model drift, dataset governance, and automated performance evaluation. Expertise in managing complex build environments (e.g., CMake, Bazel) or cloud-native infrastructure-as-code (Terraform, AWS/GCP). Direct experience navigating FDA Q-Sub or MHRA pre-submission processes. Experience with hardware-in-the-loop or software-in-the-loop test environments. Benefits: Company equity plan so all employees share in the success of the company Salary-sacrifice pension scheme Private medical, dental and vision insurance (medical history disregarded) Group life assurance at 4x annual income Comprehensive mental health support, including unlimited access to 1:1 sessions with trained professionals Unlimited holiday allowance (+ bank holidays) and one week of remote working per quarter Lunch voucher (£10) every day for JustEat and free dinner on those days where you need to work later Twice weekly deliveries of fresh fruit and an extensive selection of snacks and drinks YuLife subscription, allowing you to turn your daily steps and meditation into discounts at a range of stores
At Wayve we're committed to creating a diverse, fair and respectful culture that is inclusive of everyone based on their unique skills and perspectives, and regardless of sex, race, religion or belief, ethnic or national origin, disability, age, citizenship, marital, domestic or civil partnership status, sexual orientation, gender identity, veteran status, pregnancy or related condition (including breastfeeding) or any other basis as protected by applicable law. About us Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing the usability and safety of automated driving systems. Our vision is to create autonomy that propels the world forward. Our intelligent, mapless, and hardware-agnostic AI products are designed for automakers, accelerating the transition from assisted to automated driving. In our fast-paced environment big problems ignite us-we embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future. At Wayve, your contributions matter. We value diversity, embrace new perspectives, and foster an inclusive work environment; we back each other to deliver impact. Make Wayve the experience that defines your career! The role As a Cloud Site Reliability Engineer at Wayve, you will build and scale the reliability foundations of our AI cloud platform. This includes our Model Development Platform (powering end-to-end model development from raw data to on-road experimentation) and our GPU Compute platform (large-scale, multi-tenant GPU fleets and scheduling systems driving model training and inference at scale). This is a founding Cloud SRE role. You won't inherit a mature SRE function, you'll help create it. You will define the frameworks, automation, and operational standards that ensure our model development infrastructure, distributed systems, and large compute clusters operate predictably, efficiently, and at scale. This role sits at the intersection of AI research, large-scale cloud infrastructure, and production operations. Your work will directly enable faster model training, reliable experimentation, and scalable AI deployment by ensuring our cloud infrastructure is resilient and performant. Key responsibilities Reliability & Platform Ownership Own the reliability, availability, and performance of the Model Dev Platform and GPU Compute environments. Define and operationalise SLOs, SLIs, and error budgets across platform services. Improve capacity planning, scaling strategies, and resource efficiency across large GPU-backed clusters. Partner with ML, platform, and software teams to establish clear production readiness standards. Incident Response & On-Call Participate in a 24/7 on-call rotation as first-line response for cloud and cluster-related incidents. Lead incident triage, escalation, communications, and root cause analysis. Translate post-incident learning into durable architectural or automation improvements. Continuously reduce alert noise and recurring operational burden. Observability & Operational Excellence Design and operate monitoring, logging, tracing, and alerting systems that enable rapid detection and recovery. Build dashboards that reflect real user-centric platform health (not just infrastructure metrics). Improve deployment safety through better change management, validation, and rollback mechanisms. Automation & Tooling Build automation for cluster operations, training workflows, remediation, and scaling tasks. Implement self healing patterns and resilient recovery workflows. Harden CI/CD and release processes to improve deployment safety and velocity. Support infrastructure as code and policy driven guardrails to ensure secure, reliable cloud environments. About you In order to set you up for success as a Cloud Site Reliability Engineer at Wayve, we're looking for the following skills and experience. Essential skills Proven experience in an SRE, Production Engineer, or Cloud Reliability role supporting large-scale cloud systems. Strong Kubernetes experience, including operating production clusters. Hands on experience running production workloads in AWS, GCP, or Azure. Experience operating complex distributed systems in production, ideally including compute-heavy or high-performance workloads. Experience working with large compute clusters; exposure to AI/ML training or inference workloads strongly preferred. Strong Linux fundamentals and proficiency in at least one scripting or systems language (e.g., Python, Go, C++) with a bias toward automation. Deep troubleshooting skills across networking, storage, distributed systems, and performance at scale. Experience designing and operating observability stacks (e.g., Datadog, Prometheus, Grafana, OpenTelemetry). Clear communication skills, including leading incidents, writing post mortems, and influencing teams to prioritise reliability improvements. Desirable skills Experience operating GPU backed environments or large scale ML infrastructure. Experience running model training or inference pipelines in production (MLOps). Familiarity with infrastructure as code (e.g., Terraform) and secure cloud production environments. Experience defining and running SLOs/SLIs and building reliability programs across multiple teams. Experience as an early or founding SRE hire establishing processes from scratch. Interest in helping shape and grow a Cloud SRE function, with potential to take on leadership responsibilities over time. This is a full time role based in our office in London (2 days a week in the office). At Wayve we want the best of all worlds so we operate a hybrid working policy that combines time together in our offices and workshops to fuel innovation, culture, relationships and learning, and time spent working from home. Wayve is committed to creating an inclusive interview experience. If you require any accommodations or adjustments to participate fully in our interview process, please let us know.
22/07/2026
Full time
At Wayve we're committed to creating a diverse, fair and respectful culture that is inclusive of everyone based on their unique skills and perspectives, and regardless of sex, race, religion or belief, ethnic or national origin, disability, age, citizenship, marital, domestic or civil partnership status, sexual orientation, gender identity, veteran status, pregnancy or related condition (including breastfeeding) or any other basis as protected by applicable law. About us Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing the usability and safety of automated driving systems. Our vision is to create autonomy that propels the world forward. Our intelligent, mapless, and hardware-agnostic AI products are designed for automakers, accelerating the transition from assisted to automated driving. In our fast-paced environment big problems ignite us-we embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future. At Wayve, your contributions matter. We value diversity, embrace new perspectives, and foster an inclusive work environment; we back each other to deliver impact. Make Wayve the experience that defines your career! The role As a Cloud Site Reliability Engineer at Wayve, you will build and scale the reliability foundations of our AI cloud platform. This includes our Model Development Platform (powering end-to-end model development from raw data to on-road experimentation) and our GPU Compute platform (large-scale, multi-tenant GPU fleets and scheduling systems driving model training and inference at scale). This is a founding Cloud SRE role. You won't inherit a mature SRE function, you'll help create it. You will define the frameworks, automation, and operational standards that ensure our model development infrastructure, distributed systems, and large compute clusters operate predictably, efficiently, and at scale. This role sits at the intersection of AI research, large-scale cloud infrastructure, and production operations. Your work will directly enable faster model training, reliable experimentation, and scalable AI deployment by ensuring our cloud infrastructure is resilient and performant. Key responsibilities Reliability & Platform Ownership Own the reliability, availability, and performance of the Model Dev Platform and GPU Compute environments. Define and operationalise SLOs, SLIs, and error budgets across platform services. Improve capacity planning, scaling strategies, and resource efficiency across large GPU-backed clusters. Partner with ML, platform, and software teams to establish clear production readiness standards. Incident Response & On-Call Participate in a 24/7 on-call rotation as first-line response for cloud and cluster-related incidents. Lead incident triage, escalation, communications, and root cause analysis. Translate post-incident learning into durable architectural or automation improvements. Continuously reduce alert noise and recurring operational burden. Observability & Operational Excellence Design and operate monitoring, logging, tracing, and alerting systems that enable rapid detection and recovery. Build dashboards that reflect real user-centric platform health (not just infrastructure metrics). Improve deployment safety through better change management, validation, and rollback mechanisms. Automation & Tooling Build automation for cluster operations, training workflows, remediation, and scaling tasks. Implement self healing patterns and resilient recovery workflows. Harden CI/CD and release processes to improve deployment safety and velocity. Support infrastructure as code and policy driven guardrails to ensure secure, reliable cloud environments. About you In order to set you up for success as a Cloud Site Reliability Engineer at Wayve, we're looking for the following skills and experience. Essential skills Proven experience in an SRE, Production Engineer, or Cloud Reliability role supporting large-scale cloud systems. Strong Kubernetes experience, including operating production clusters. Hands on experience running production workloads in AWS, GCP, or Azure. Experience operating complex distributed systems in production, ideally including compute-heavy or high-performance workloads. Experience working with large compute clusters; exposure to AI/ML training or inference workloads strongly preferred. Strong Linux fundamentals and proficiency in at least one scripting or systems language (e.g., Python, Go, C++) with a bias toward automation. Deep troubleshooting skills across networking, storage, distributed systems, and performance at scale. Experience designing and operating observability stacks (e.g., Datadog, Prometheus, Grafana, OpenTelemetry). Clear communication skills, including leading incidents, writing post mortems, and influencing teams to prioritise reliability improvements. Desirable skills Experience operating GPU backed environments or large scale ML infrastructure. Experience running model training or inference pipelines in production (MLOps). Familiarity with infrastructure as code (e.g., Terraform) and secure cloud production environments. Experience defining and running SLOs/SLIs and building reliability programs across multiple teams. Experience as an early or founding SRE hire establishing processes from scratch. Interest in helping shape and grow a Cloud SRE function, with potential to take on leadership responsibilities over time. This is a full time role based in our office in London (2 days a week in the office). At Wayve we want the best of all worlds so we operate a hybrid working policy that combines time together in our offices and workshops to fuel innovation, culture, relationships and learning, and time spent working from home. Wayve is committed to creating an inclusive interview experience. If you require any accommodations or adjustments to participate fully in our interview process, please let us know.
Thought Machine's mission is bold - to properly and permanently rid the world's banks of legacy technology. To achieve this, we have developed the foundations of modern banking through core and payments technology which run natively in the cloud. What we are attempting is hard and means we need great people working together to build great technology. We have grown rapidly in the past few years - growing our team to more than 550 individuals across offices in London, New York, Singapore and Sydney. We have raised more than $500m in funding and are now valued at $2.7bn. Our investors include Molten Ventures, Eurazeo, Intesa Sanpaolo, Temasek, Nyca Partners, JPMorgan Chase Strategic Investments, Standard Chartered Ventures, and more. We have created a culture that enables our team to produce the best work in the industry while ensuring we have fun along the way. We're regularly cited as having a fantastic workplace culture and have been recognised by Sifted magazine as having one of the highest Glassdoor ratings for a UK fintech company and the industry's most generous employee share package. Named one of the world's most innovative fintechs by Global Finance Magazine, we were also recognised by the Financial Times as one of Europe's fastest-growing companies for two consecutive years-and a UK Best Employer for 2026. As a Senior Cloud Support Engineer at Thought Machine, you're stepping into a senior technical role that goes well beyond incident response. At this level, you're expected to own, influence and elevate how we deliver support globally, across both our hosted SaaS product and bank-hosted Vault deployments. You'll not only handle our most complex technical issues, but also mentor junior engineers, shape tooling and processes, partnering closely with Engineering and Product teams to drive continuous improvement across the platform. You bring a deep understanding of cloud-native systems and distributed architecture, and will become a domain expert in Vault Core. This is a highly visible role where your technical judgement, leadership and strategic problem-solving will directly impact client outcomes and platform quality. Duties Lead the investigation and resolution of high-priority or complex incidents, coordinating with engineering and delivery teams to drive timely solutions Own escalations end-to-end, providing technical depth, structure and calm during high-pressure client situations Develop and refine tools, dashboards and automation to improve support delivery, observability and onboarding Identify recurring issues, propose and lead solutions that improve platform stability, reduce effort and enhance client experience Provide mentoring, training and technical oversight for IC1 and IC2 engineers Contribute to internal documentation, playbooks and root cause analysis reports, with a focus on accuracy and knowledge sharing Support both hosted and self-managed deployments of Vault, collaborating with client teams on best practice configurations Own and maintain Enterprise relationships from a technical perspective Actively shape the evolution of our support model, platform tooling and processes across regions Essential 5+ years of experience in technical support, cloud infrastructure or SRE/DevOps roles, ideally within high-scale, client-facing environments Deep experience with incident management, root cause analysis and production troubleshooting Strong Linux systems knowledge, including filesystems, networking and system internals Programming skills in Golang and Python and experience with infrastructure tools or observability stacks (e.g. Grafana, Prometheus, EFK) Confidence in working with cloud-native platforms and tools (e.g. Kubernetes, Terraform, AWS/GCP, Docker) Excellent communication skills under pressure, being able to influence across client, product and engineering boundaries Demonstrated ability to lead without authority; especially in escalations, cross-functional collaborations and mentoring A drive to raise the standard for support and build systems that scale sustainably Desirable Experience deploying or supporting self-managed software in client cloud environments Exposure to core banking, payments or financial services infrastructure Familiarity with gRPC, protobuf or service-to-service communication patterns Open-source contributions or evidence of self-driven learning in cloud, infra or backend tooling Experience contributing to support process design or automation initiatives Benefits Highly competitive salary and commission Pension plan (match up to 5%) Life insurance (3x annual salary) Competitive maternity (six months fully paid) and paternity leave (four weeks fully paid) Shared parental leave (matched to our maternity leave for the same point in time) 25 days holiday and bank holidays Flexible working hours Cycle-to-work scheme Electric car scheme Season ticket loan Access to outstanding learning materials and courses Sports and hobby clubs, subsidised by Thought Machine All the latest tech you need Start the day properly with fresh fruit and cereals Huge range of healthy (and not-so-healthy) snacks, smoothies and drinks A talented and experienced team as your colleagues An environment where we encourage learning and progress Two charity days a year Weekly food pop-up We actively hire candidates who demonstrate technical excellence in their field and welcome people of all ages and backgrounds, providing everyone with equal access to professional development. You are encouraged to apply even if your experience doesn't accurately match the job description. We also encourage applications from those with different abilities, including candidates with ADHD, autism, dyslexia or dyspraxia.
22/07/2026
Full time
Thought Machine's mission is bold - to properly and permanently rid the world's banks of legacy technology. To achieve this, we have developed the foundations of modern banking through core and payments technology which run natively in the cloud. What we are attempting is hard and means we need great people working together to build great technology. We have grown rapidly in the past few years - growing our team to more than 550 individuals across offices in London, New York, Singapore and Sydney. We have raised more than $500m in funding and are now valued at $2.7bn. Our investors include Molten Ventures, Eurazeo, Intesa Sanpaolo, Temasek, Nyca Partners, JPMorgan Chase Strategic Investments, Standard Chartered Ventures, and more. We have created a culture that enables our team to produce the best work in the industry while ensuring we have fun along the way. We're regularly cited as having a fantastic workplace culture and have been recognised by Sifted magazine as having one of the highest Glassdoor ratings for a UK fintech company and the industry's most generous employee share package. Named one of the world's most innovative fintechs by Global Finance Magazine, we were also recognised by the Financial Times as one of Europe's fastest-growing companies for two consecutive years-and a UK Best Employer for 2026. As a Senior Cloud Support Engineer at Thought Machine, you're stepping into a senior technical role that goes well beyond incident response. At this level, you're expected to own, influence and elevate how we deliver support globally, across both our hosted SaaS product and bank-hosted Vault deployments. You'll not only handle our most complex technical issues, but also mentor junior engineers, shape tooling and processes, partnering closely with Engineering and Product teams to drive continuous improvement across the platform. You bring a deep understanding of cloud-native systems and distributed architecture, and will become a domain expert in Vault Core. This is a highly visible role where your technical judgement, leadership and strategic problem-solving will directly impact client outcomes and platform quality. Duties Lead the investigation and resolution of high-priority or complex incidents, coordinating with engineering and delivery teams to drive timely solutions Own escalations end-to-end, providing technical depth, structure and calm during high-pressure client situations Develop and refine tools, dashboards and automation to improve support delivery, observability and onboarding Identify recurring issues, propose and lead solutions that improve platform stability, reduce effort and enhance client experience Provide mentoring, training and technical oversight for IC1 and IC2 engineers Contribute to internal documentation, playbooks and root cause analysis reports, with a focus on accuracy and knowledge sharing Support both hosted and self-managed deployments of Vault, collaborating with client teams on best practice configurations Own and maintain Enterprise relationships from a technical perspective Actively shape the evolution of our support model, platform tooling and processes across regions Essential 5+ years of experience in technical support, cloud infrastructure or SRE/DevOps roles, ideally within high-scale, client-facing environments Deep experience with incident management, root cause analysis and production troubleshooting Strong Linux systems knowledge, including filesystems, networking and system internals Programming skills in Golang and Python and experience with infrastructure tools or observability stacks (e.g. Grafana, Prometheus, EFK) Confidence in working with cloud-native platforms and tools (e.g. Kubernetes, Terraform, AWS/GCP, Docker) Excellent communication skills under pressure, being able to influence across client, product and engineering boundaries Demonstrated ability to lead without authority; especially in escalations, cross-functional collaborations and mentoring A drive to raise the standard for support and build systems that scale sustainably Desirable Experience deploying or supporting self-managed software in client cloud environments Exposure to core banking, payments or financial services infrastructure Familiarity with gRPC, protobuf or service-to-service communication patterns Open-source contributions or evidence of self-driven learning in cloud, infra or backend tooling Experience contributing to support process design or automation initiatives Benefits Highly competitive salary and commission Pension plan (match up to 5%) Life insurance (3x annual salary) Competitive maternity (six months fully paid) and paternity leave (four weeks fully paid) Shared parental leave (matched to our maternity leave for the same point in time) 25 days holiday and bank holidays Flexible working hours Cycle-to-work scheme Electric car scheme Season ticket loan Access to outstanding learning materials and courses Sports and hobby clubs, subsidised by Thought Machine All the latest tech you need Start the day properly with fresh fruit and cereals Huge range of healthy (and not-so-healthy) snacks, smoothies and drinks A talented and experienced team as your colleagues An environment where we encourage learning and progress Two charity days a year Weekly food pop-up We actively hire candidates who demonstrate technical excellence in their field and welcome people of all ages and backgrounds, providing everyone with equal access to professional development. You are encouraged to apply even if your experience doesn't accurately match the job description. We also encourage applications from those with different abilities, including candidates with ADHD, autism, dyslexia or dyspraxia.
At CoMind, we are developing a non-invasive neuromonitoring technology that will result in a new era of clinical brain monitoring. In joining us, you will be helping to create cutting-edge technologies that will improve how we diagnose and treat brain disorders, ultimately improving and saving the lives of patients across the world. The Role: The Software Engineering team at CoMind spans the full stack of a regulated medical device company from embedded and real-time systems through application software, ML inference, GUI, DevOps, and test engineering. It is a team that builds across every layer, working to medical device software standards and shipping software that has to perform reliably in clinical environments. As a Senior/Staff Software Test Engineer, you will define the quality and verification strategy across CoMind's entire software stack from embedded real-time systems to cloud-side data pipelines. You will move beyond task-level execution to architect the infrastructure and 'compliance-as-code' processes that provide the engineering team with high-velocity, high-confidence validation. You will serve as a key technical authority, ensuring our software not only meets stringent IEC 62304 and ISO 14971 standards but does so through sustainable, automated design. You will function as the technical authority for software quality. Your goal is to move CoMind beyond traditional V&V, architecting a self-validating system that integrates safety, regulatory compliance, and high-velocity development. You will be the bridge between Systems Engineering, Regulatory Affairs, and Software Engineering, ensuring that quality is an inherent attribute of our technology rather than a post-development checklist. At CoMind, all team members work at least 4 days per week in the office, plus a flexible work-from-home day. This role is based in our London (Kings Cross) office. Responsibilities: Architect a scalable, high-fidelity testing infrastructure that enables continuous validation. Champion 'design for testability' across the entire engineering stack, influencing system architecture to ensure critical safety and performance requirements are verifiable by design. Define the end-to-end V&V strategy for the product lifecycle. Quantitatively map software verification evidence directly to system-level safety risks (ISO 14971) to create a 'compliance-as-code' environment that drastically reduces time-to-submission for FDA and other regulatory bodies. Drive a 'quality-first' engineering culture by mentorship. Lead training and initiatives on Test-Driven Development (TDD), CI/CD best practices, and automated quality gates, ensuring that the engineering team's velocity is sustained by high-confidence automated processes. Lead the development of AI-augmented test strategies. Architect predictive anomaly detection systems and AI-driven test generation engines that evolve alongside our software, turning the test suite into an intelligent system that learns from past regressions. Collaborate with engineering leadership to drive continuous improvement in software development and test processes, ensuring that quality practices scale with the organisation. AI is fundamental to our culture - it's not just a tool, but a core part of how we work, collaborate, and innovate. We expect all team members to embrace AI in their daily work and continuously find new ways to use it effectively. Skills & Experience: Strategic Execution: 7+ years in software test engineering, with a proven track record of architecting testing strategies for complex, regulated systems (embedded, real-time, or SaMD). Regulatory & Risk Authority: Deep working knowledge of IEC 62304 and ISO 14971. Proven ability to design 'compliance-as-code' workflows that integrate risk controls directly into the CI/CD pipeline and automate regulatory evidence generation. Technical Proficiency: Advanced Python and C/C++; extensive experience building scalable test infrastructure for both embedded hardware and cloud-side data pipelines. Cultural Leadership: Demonstrated success in mentoring software engineers and driving a 'quality-first' culture through TDD, static analysis, and automated quality gates. Cybersecurity & Documentation: Strong experience in cybersecurity requirements (e.g., IEC 81001-5-1, FDA cybersecurity guidance) and rigorous documentation management for regulatory submission. Nice to Have Experience testing real-time Linux or RTOS-based products. Background in MLOps, including validation of model drift, dataset governance, and automated performance evaluation. Expertise in managing complex build environments (e.g., CMake, Bazel) or cloud-native infrastructure-as-code (Terraform, AWS/GCP). Direct experience navigating FDA Q-Sub or MHRA pre-submission processes. Experience with hardware-in-the-loop or software-in-the-loop test environments. Benefits: Company equity plan so all employees share in the success of the company Salary-sacrifice pension scheme Private medical, dental and vision insurance (medical history disregarded) Group life assurance at 4x annual income Comprehensive mental health support, including unlimited access to 1:1 sessions with trained professionals Unlimited holiday allowance (+ bank holidays) and one week of remote working per quarter Lunch voucher (£10) every day for JustEat and free dinner on those days where you need to work later Twice weekly deliveries of fresh fruit and an extensive selection of snacks and drinks YuLife subscription, allowing you to turn your daily steps and meditation into discounts at a range of stores
19/07/2026
Full time
At CoMind, we are developing a non-invasive neuromonitoring technology that will result in a new era of clinical brain monitoring. In joining us, you will be helping to create cutting-edge technologies that will improve how we diagnose and treat brain disorders, ultimately improving and saving the lives of patients across the world. The Role: The Software Engineering team at CoMind spans the full stack of a regulated medical device company from embedded and real-time systems through application software, ML inference, GUI, DevOps, and test engineering. It is a team that builds across every layer, working to medical device software standards and shipping software that has to perform reliably in clinical environments. As a Senior/Staff Software Test Engineer, you will define the quality and verification strategy across CoMind's entire software stack from embedded real-time systems to cloud-side data pipelines. You will move beyond task-level execution to architect the infrastructure and 'compliance-as-code' processes that provide the engineering team with high-velocity, high-confidence validation. You will serve as a key technical authority, ensuring our software not only meets stringent IEC 62304 and ISO 14971 standards but does so through sustainable, automated design. You will function as the technical authority for software quality. Your goal is to move CoMind beyond traditional V&V, architecting a self-validating system that integrates safety, regulatory compliance, and high-velocity development. You will be the bridge between Systems Engineering, Regulatory Affairs, and Software Engineering, ensuring that quality is an inherent attribute of our technology rather than a post-development checklist. At CoMind, all team members work at least 4 days per week in the office, plus a flexible work-from-home day. This role is based in our London (Kings Cross) office. Responsibilities: Architect a scalable, high-fidelity testing infrastructure that enables continuous validation. Champion 'design for testability' across the entire engineering stack, influencing system architecture to ensure critical safety and performance requirements are verifiable by design. Define the end-to-end V&V strategy for the product lifecycle. Quantitatively map software verification evidence directly to system-level safety risks (ISO 14971) to create a 'compliance-as-code' environment that drastically reduces time-to-submission for FDA and other regulatory bodies. Drive a 'quality-first' engineering culture by mentorship. Lead training and initiatives on Test-Driven Development (TDD), CI/CD best practices, and automated quality gates, ensuring that the engineering team's velocity is sustained by high-confidence automated processes. Lead the development of AI-augmented test strategies. Architect predictive anomaly detection systems and AI-driven test generation engines that evolve alongside our software, turning the test suite into an intelligent system that learns from past regressions. Collaborate with engineering leadership to drive continuous improvement in software development and test processes, ensuring that quality practices scale with the organisation. AI is fundamental to our culture - it's not just a tool, but a core part of how we work, collaborate, and innovate. We expect all team members to embrace AI in their daily work and continuously find new ways to use it effectively. Skills & Experience: Strategic Execution: 7+ years in software test engineering, with a proven track record of architecting testing strategies for complex, regulated systems (embedded, real-time, or SaMD). Regulatory & Risk Authority: Deep working knowledge of IEC 62304 and ISO 14971. Proven ability to design 'compliance-as-code' workflows that integrate risk controls directly into the CI/CD pipeline and automate regulatory evidence generation. Technical Proficiency: Advanced Python and C/C++; extensive experience building scalable test infrastructure for both embedded hardware and cloud-side data pipelines. Cultural Leadership: Demonstrated success in mentoring software engineers and driving a 'quality-first' culture through TDD, static analysis, and automated quality gates. Cybersecurity & Documentation: Strong experience in cybersecurity requirements (e.g., IEC 81001-5-1, FDA cybersecurity guidance) and rigorous documentation management for regulatory submission. Nice to Have Experience testing real-time Linux or RTOS-based products. Background in MLOps, including validation of model drift, dataset governance, and automated performance evaluation. Expertise in managing complex build environments (e.g., CMake, Bazel) or cloud-native infrastructure-as-code (Terraform, AWS/GCP). Direct experience navigating FDA Q-Sub or MHRA pre-submission processes. Experience with hardware-in-the-loop or software-in-the-loop test environments. Benefits: Company equity plan so all employees share in the success of the company Salary-sacrifice pension scheme Private medical, dental and vision insurance (medical history disregarded) Group life assurance at 4x annual income Comprehensive mental health support, including unlimited access to 1:1 sessions with trained professionals Unlimited holiday allowance (+ bank holidays) and one week of remote working per quarter Lunch voucher (£10) every day for JustEat and free dinner on those days where you need to work later Twice weekly deliveries of fresh fruit and an extensive selection of snacks and drinks YuLife subscription, allowing you to turn your daily steps and meditation into discounts at a range of stores
Responsibilities and Impact Own and execute the AIOps product roadmap, aligning priorities with enterprise reliability goals, operational maturity, adoption targets, and measurable business outcomes. Translate operational needs into clear product requirements, use cases, user stories, acceptance criteria, and prioritized backlog items across the AIOps product lifecycle. Ownership of high-impact AIOps capabilities, including noise reduction, event management, correlation, anomaly detection, root cause analysis, predictive alerting, and self healing remediation. Partner across Operations, infrastructure, SRE, platform engineering, service management, application teams, security, and vendors to operationalize AIOps capabilities and drive adoption at enterprise scale. Define and track success metrics for alert quality, incident reduction, MTTD, MTTR, service health, automation coverage, platform adoption, and operational efficiency. Present roadmap progress, risks, decisions, and outcome metrics to senior stakeholders to influence improvements to tooling, workflow, and the operating model. Maintain an analytical and data driven mindset with strong problem solving, prioritization, and decision making skills. Basic Required Qualifications 10+ years of experience in product management, IT operations, SRE, observability, platform engineering, or related enterprise technology roles. Strong understanding of AIOps concepts, including event correlation, anomaly detection, root cause analysis, noise reduction, predictive analytics, and automated remediation. Experience defining product roadmaps, managing requirements and backlog priorities, and delivering technical platform capabilities across the product lifecycle. Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent practical experience. Machine Learning: Expertise in time series forecasting, clustering, and correlation algorithms. Generative AI: Experience with LLM prompting, Retrieval Augmented Generation (RAG), and LangChain. Cloud Infrastructure: Deep understanding of AWS, Azure, or GCP infrastructure components. Additional Preferred Qualifications Experience with enterprise observability and AIOps platforms such as ServiceNow ITOM, Splunk, Dynatrace, Datadog, BigPanda, New Relic, or similar platforms. Familiarity with cloud native platforms, container orchestration systems, distributed systems, service mapping, and telemetry standards. Experience defining KPIs, SLOs, product adoption measures, and operational maturity frameworks for enterprise reliability and service performance. Relevant certifications in product management, cloud, ITSM, observability, Agile, or platform engineering are a plus. Benefits Health & Wellness: Health care coverage designed for the mind and body. Flexible Downtime: Generous time off helps keep you energized for your time on. Continuous Learning: Access a wealth of resources to grow your career and learn valuable new skills. Invest in Your Future: Secure your financial future through competitive pay, retirement planning, a continuing education program with a company matched student loan contribution, and financial wellness programs. Family Friendly Perks: It's not just about you. S&P Global has perks for your partners and little ones, too, with some best in class benefits for families. Beyond the Basics: From retail discounts to referral incentive awards-small perks can make a big difference. Equal Opportunity Employer S&P Global is an equal opportunity employer and all qualified candidates will receive consideration for employment without regard to race/ethnicity, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, marital status, military veteran status, unemployment status, or any other status protected by law. Only electronic job submissions will be considered for employment. If you need an accommodation during the application process due to a disability, please send an email to: and your request will be forwarded to the appropriate person.
19/07/2026
Full time
Responsibilities and Impact Own and execute the AIOps product roadmap, aligning priorities with enterprise reliability goals, operational maturity, adoption targets, and measurable business outcomes. Translate operational needs into clear product requirements, use cases, user stories, acceptance criteria, and prioritized backlog items across the AIOps product lifecycle. Ownership of high-impact AIOps capabilities, including noise reduction, event management, correlation, anomaly detection, root cause analysis, predictive alerting, and self healing remediation. Partner across Operations, infrastructure, SRE, platform engineering, service management, application teams, security, and vendors to operationalize AIOps capabilities and drive adoption at enterprise scale. Define and track success metrics for alert quality, incident reduction, MTTD, MTTR, service health, automation coverage, platform adoption, and operational efficiency. Present roadmap progress, risks, decisions, and outcome metrics to senior stakeholders to influence improvements to tooling, workflow, and the operating model. Maintain an analytical and data driven mindset with strong problem solving, prioritization, and decision making skills. Basic Required Qualifications 10+ years of experience in product management, IT operations, SRE, observability, platform engineering, or related enterprise technology roles. Strong understanding of AIOps concepts, including event correlation, anomaly detection, root cause analysis, noise reduction, predictive analytics, and automated remediation. Experience defining product roadmaps, managing requirements and backlog priorities, and delivering technical platform capabilities across the product lifecycle. Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent practical experience. Machine Learning: Expertise in time series forecasting, clustering, and correlation algorithms. Generative AI: Experience with LLM prompting, Retrieval Augmented Generation (RAG), and LangChain. Cloud Infrastructure: Deep understanding of AWS, Azure, or GCP infrastructure components. Additional Preferred Qualifications Experience with enterprise observability and AIOps platforms such as ServiceNow ITOM, Splunk, Dynatrace, Datadog, BigPanda, New Relic, or similar platforms. Familiarity with cloud native platforms, container orchestration systems, distributed systems, service mapping, and telemetry standards. Experience defining KPIs, SLOs, product adoption measures, and operational maturity frameworks for enterprise reliability and service performance. Relevant certifications in product management, cloud, ITSM, observability, Agile, or platform engineering are a plus. Benefits Health & Wellness: Health care coverage designed for the mind and body. Flexible Downtime: Generous time off helps keep you energized for your time on. Continuous Learning: Access a wealth of resources to grow your career and learn valuable new skills. Invest in Your Future: Secure your financial future through competitive pay, retirement planning, a continuing education program with a company matched student loan contribution, and financial wellness programs. Family Friendly Perks: It's not just about you. S&P Global has perks for your partners and little ones, too, with some best in class benefits for families. Beyond the Basics: From retail discounts to referral incentive awards-small perks can make a big difference. Equal Opportunity Employer S&P Global is an equal opportunity employer and all qualified candidates will receive consideration for employment without regard to race/ethnicity, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, marital status, military veteran status, unemployment status, or any other status protected by law. Only electronic job submissions will be considered for employment. If you need an accommodation during the application process due to a disability, please send an email to: and your request will be forwarded to the appropriate person.
Music is Universal It's the passionate and dedicated team at Universal Music who help make us the world's leading music company. From A&R to finance, legal to digital, sales to marketing, Universal Music is the place to grow and develop your career within a truly commercial and innovative business that leads in everything it does.Everyone is welcome to apply for our roles, and we are determined to ensure that no applicant or employee receives less favourable treatment because of gender, race, disability, sexual orientation, religion, belief, age, marital status, background, pregnancy, or caring responsibilities. We also recognise the importance of diversity of thought within our teams and are fully committed to embracing the talents of people with autism, dyslexia, ADHD, and other forms of neurocognitive variation.We will always seek to make appropriate adjustments to recruitment, workplaces, and work processes to be fully inclusive to people with different needs and working styles. If you need us to make any reasonable adjustments for you from application onwards, including alternatives to the online form or to disclose a neurocognitive condition, please email . Job Summary: We are UMG, the Universal Music Group. We are the world's leading music company. In everything we do, we are committed to artistry, innovation and entrepreneurship. We own and operate a broad array of businesses engaged in recorded music, music publishing, merchandising, and audiovisual content in more than 60 countries. We identify and develop recording artists and songwriters, and we produce, distribute and promote the most critically acclaimed and commercially successful music to delight and entertain fans around the world.As a Senior Observability Engineer, you will be a driving force for technical excellence and strategic vision within our global team. You will be instrumental in architecting, building, and leading our comprehensive observability strategy to ensure the reliability, performance, and scalability of our critical IT systems. This senior role demands a passion for data-driven strategy, a commitment to automation, and the ability to mentor and lead. You will not only solve complex technical challenges but also influence the direction of observability practices across UMG globally, ensuring our technology landscape is as world-class as our music. Job Functions: Architecture & Strategy: Lead the architectural design and strategic roadmap for our observability stack. Drive the vision for world-class monitoring, logging, tracing, and alerting solutions across our hybrid and cloud-native environments. Innovate & Automate: Spearhead the evaluation, selection, and implementation of cutting-edge observability tools and platforms (e.g., Dynatrace, OpenTelemetry, Prometheus, Grafana). Architect and build robust, automated observability pipelines. Take an active part in documenting and defining processes and best practice. Optimize & Analyze: Conduct deep-dive analysis of telemetry data to proactively identify performance bottlenecks, optimize resource utilization, and guide capacity planning. Lead & Mentor: Act as a technical leader and mentor for the observability team and wider engineering groups. Champion and enforce best practices, fostering a culture of proactive and data-informed decision-making. Drive Incident & Problem Management: Working with Operations teams on high-priority incident resolution efforts, utilizing deep analysis of telemetry data for swift root cause identification. Drive post-incident reviews and implement long-term solutions to enhance system resilience. Collaborate & Influence: Partner with Development, SRE, and Infrastructure leaders to embed observability into the entire technology lifecycle. Influence and drive the adoption of observability best practices across the global organization. Champion the use of observability in the global UMG environment. Make UMG the place to be: Mentoring, managing and genuinely leading the Observability team in a way that attracts and retains the best talent. UMG is a place where everyone can bring themselves fully to work and thrive, as a Leader you are a key part of this. Job Requirements: Essential Qualifications Experience: 5-7+ years of hands-on experience in an Observability, Site Reliability Engineering (SRE), or DevOps role, with a proven track record of leading complex projects. Technical Leadership: Demonstrated experience in architecting and designing large-scale monitoring and observability solutions. Expert-Level Tooling: Deep expertise with modern observability platforms (e.g., Dynatrace, AWS Cloudwatch, Prometheus, Grafana, ELK Stack, Splunk, OpenTelemetry). Cloud & Infrastructure: Advanced knowledge of major cloud platforms (AWS, Azure, GCP), containerization (Docker, Kubernetes), and Infrastructure as Code (Terraform, Ansible). Programming & Automation: Strong programming and scripting skills (e.g., Python, Go, Shell) with a focus on creating scalable automation and custom tooling. Problem-Solving: Exceptional analytical and strategic problem-solving skills, with the ability to lead through complex technical challenges. Data Analysis: Expertise in analysing and visualising telemetry data into meaningful information to drive actions. Hands-on: Demonstratable hands-on engineering and coding experience, ability to deep-dive into existing and emerging technologies to identify opportunities and solutions. Containerization and Orchestration: Understanding of container technologies (e.g., Docker) and container orchestration platforms (e.g., Kubernetes) to monitor and manage containerized applications. Networking Knowledge: Understanding of networking principles and protocols to effectively monitor and troubleshoot network-related issues. Security Awareness: Awareness of security best practices and the ability to integrate security monitoring into observability processes. Communication & Influence: Excellent communication and interpersonal skills, capable of articulating a technical vision to diverse audiences and influencing senior stakeholders. Ability to collaborate with cross-functional teams, convey findings, and discuss improvements with developers and operations teams. Continuous Learning: Given the dynamic nature of technology, a commitment to continuous learning and staying updated on the latest trends in observability and monitoring. Self-motivated with a high degree of initiative and excellent follow-up skills, along with strong analytical and problem-solving skills. Travel may be required but is not part of the regular work schedule. Bachelor's degree in technology related field as well as 5+ years of relevant experience within the Observability field.Desired Qualifications Advanced Concepts: Proven experience with Chaos Engineering, AI-driven analytics, defining SLOs/SLIs, and advanced deployment strategies (Canary/Blue-Green). Software Engineering Foundation: Strong background in software engineering principles, database administration, and distributed systems architecture Certifications: Relevant senior-level industry certifications (e.g., AWS Certified DevOps Engineer - Professional, Certified Kubernetes Administrator).Just So You Know The company presents this job description as a guide to the major areas and duties for which the jobholder is accountable. However, the business operates in an environment that demands change and the jobholder's specific responsibilities and activities will vary and develop. Therefore, the job description should be seen as indicative and not as a permanent, definitive, and exhaustive statement. Job Category: Universal Music Group
18/07/2026
Full time
Music is Universal It's the passionate and dedicated team at Universal Music who help make us the world's leading music company. From A&R to finance, legal to digital, sales to marketing, Universal Music is the place to grow and develop your career within a truly commercial and innovative business that leads in everything it does.Everyone is welcome to apply for our roles, and we are determined to ensure that no applicant or employee receives less favourable treatment because of gender, race, disability, sexual orientation, religion, belief, age, marital status, background, pregnancy, or caring responsibilities. We also recognise the importance of diversity of thought within our teams and are fully committed to embracing the talents of people with autism, dyslexia, ADHD, and other forms of neurocognitive variation.We will always seek to make appropriate adjustments to recruitment, workplaces, and work processes to be fully inclusive to people with different needs and working styles. If you need us to make any reasonable adjustments for you from application onwards, including alternatives to the online form or to disclose a neurocognitive condition, please email . Job Summary: We are UMG, the Universal Music Group. We are the world's leading music company. In everything we do, we are committed to artistry, innovation and entrepreneurship. We own and operate a broad array of businesses engaged in recorded music, music publishing, merchandising, and audiovisual content in more than 60 countries. We identify and develop recording artists and songwriters, and we produce, distribute and promote the most critically acclaimed and commercially successful music to delight and entertain fans around the world.As a Senior Observability Engineer, you will be a driving force for technical excellence and strategic vision within our global team. You will be instrumental in architecting, building, and leading our comprehensive observability strategy to ensure the reliability, performance, and scalability of our critical IT systems. This senior role demands a passion for data-driven strategy, a commitment to automation, and the ability to mentor and lead. You will not only solve complex technical challenges but also influence the direction of observability practices across UMG globally, ensuring our technology landscape is as world-class as our music. Job Functions: Architecture & Strategy: Lead the architectural design and strategic roadmap for our observability stack. Drive the vision for world-class monitoring, logging, tracing, and alerting solutions across our hybrid and cloud-native environments. Innovate & Automate: Spearhead the evaluation, selection, and implementation of cutting-edge observability tools and platforms (e.g., Dynatrace, OpenTelemetry, Prometheus, Grafana). Architect and build robust, automated observability pipelines. Take an active part in documenting and defining processes and best practice. Optimize & Analyze: Conduct deep-dive analysis of telemetry data to proactively identify performance bottlenecks, optimize resource utilization, and guide capacity planning. Lead & Mentor: Act as a technical leader and mentor for the observability team and wider engineering groups. Champion and enforce best practices, fostering a culture of proactive and data-informed decision-making. Drive Incident & Problem Management: Working with Operations teams on high-priority incident resolution efforts, utilizing deep analysis of telemetry data for swift root cause identification. Drive post-incident reviews and implement long-term solutions to enhance system resilience. Collaborate & Influence: Partner with Development, SRE, and Infrastructure leaders to embed observability into the entire technology lifecycle. Influence and drive the adoption of observability best practices across the global organization. Champion the use of observability in the global UMG environment. Make UMG the place to be: Mentoring, managing and genuinely leading the Observability team in a way that attracts and retains the best talent. UMG is a place where everyone can bring themselves fully to work and thrive, as a Leader you are a key part of this. Job Requirements: Essential Qualifications Experience: 5-7+ years of hands-on experience in an Observability, Site Reliability Engineering (SRE), or DevOps role, with a proven track record of leading complex projects. Technical Leadership: Demonstrated experience in architecting and designing large-scale monitoring and observability solutions. Expert-Level Tooling: Deep expertise with modern observability platforms (e.g., Dynatrace, AWS Cloudwatch, Prometheus, Grafana, ELK Stack, Splunk, OpenTelemetry). Cloud & Infrastructure: Advanced knowledge of major cloud platforms (AWS, Azure, GCP), containerization (Docker, Kubernetes), and Infrastructure as Code (Terraform, Ansible). Programming & Automation: Strong programming and scripting skills (e.g., Python, Go, Shell) with a focus on creating scalable automation and custom tooling. Problem-Solving: Exceptional analytical and strategic problem-solving skills, with the ability to lead through complex technical challenges. Data Analysis: Expertise in analysing and visualising telemetry data into meaningful information to drive actions. Hands-on: Demonstratable hands-on engineering and coding experience, ability to deep-dive into existing and emerging technologies to identify opportunities and solutions. Containerization and Orchestration: Understanding of container technologies (e.g., Docker) and container orchestration platforms (e.g., Kubernetes) to monitor and manage containerized applications. Networking Knowledge: Understanding of networking principles and protocols to effectively monitor and troubleshoot network-related issues. Security Awareness: Awareness of security best practices and the ability to integrate security monitoring into observability processes. Communication & Influence: Excellent communication and interpersonal skills, capable of articulating a technical vision to diverse audiences and influencing senior stakeholders. Ability to collaborate with cross-functional teams, convey findings, and discuss improvements with developers and operations teams. Continuous Learning: Given the dynamic nature of technology, a commitment to continuous learning and staying updated on the latest trends in observability and monitoring. Self-motivated with a high degree of initiative and excellent follow-up skills, along with strong analytical and problem-solving skills. Travel may be required but is not part of the regular work schedule. Bachelor's degree in technology related field as well as 5+ years of relevant experience within the Observability field.Desired Qualifications Advanced Concepts: Proven experience with Chaos Engineering, AI-driven analytics, defining SLOs/SLIs, and advanced deployment strategies (Canary/Blue-Green). Software Engineering Foundation: Strong background in software engineering principles, database administration, and distributed systems architecture Certifications: Relevant senior-level industry certifications (e.g., AWS Certified DevOps Engineer - Professional, Certified Kubernetes Administrator).Just So You Know The company presents this job description as a guide to the major areas and duties for which the jobholder is accountable. However, the business operates in an environment that demands change and the jobholder's specific responsibilities and activities will vary and develop. Therefore, the job description should be seen as indicative and not as a permanent, definitive, and exhaustive statement. Job Category: Universal Music Group
About aion Aion is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, Aion takes them from concept to production on a single, unified platform. We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters. We're a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we're building the next generation of enterprise AI and we're looking for exceptional people to help us scale. Who You Are You're a hands-on AI engineer with 3-5+ years of experience building production-grade multimodal AI systems and LLM applications. Your responsibilities mirror those of a hands-on AI startup CTO you work in small teams to own delivery of high-stakes customer projects, embedding directly at client sites to architect, build, and deploy intelligent agent solutions. You're equally comfortable writing production code, presenting technical solutions to C-level executives, and debugging complex AI systems on factory floors or in customer data centers. You've shipped voice agents, video processing systems, or conversational AI to production. You thrive translating ambiguous business requirements into concrete technical solutions that create measurable impact. You're comfortable working across the full AI deployment lifecycle from use case discovery and solution architecture to multimodal agent development, MLOps pipeline implementation, and production optimization. You understand what makes agents perform well in production and how to systematically improve quality through observability and evaluation. Experience with voice AI platforms, RAG systems, and LLM orchestration frameworks is highly desirable. You bring exceptional communication skills, customer empathy, and the drive to build AI solutions that transform enterprise operations globally. What You'll Do Customer Engagement & Multimodal Agent Development Work directly at customer sites from factory floors to executive offices conducting discovery workshops and technical assessments to identify high-impact AI opportunities Design and architect end-to-end multimodal agent systems (voice + video + text) that leverage aion's distributed GPU infrastructure and managed services Build production-grade voice AI systems using STT, TTS APIs, and LLMs deployed on aion's platform Develop vision-enabled agents processing real-time video streams using computer vision pipelines on aion's infrastructure Implement sophisticated multi-agent orchestration with(or similar) frameworks like LangChain or LlamaIndex-enabling tool use, memory management, and autonomous task completion Rapidly prototype POCs in 2-4 weeks, coding alongside client teams to validate concepts and iterate based on feedback Optimize for sub-500ms latency, natural conversation flow, turn detection, and interruption handling in real-time systems Integrate agents directly into customer codebases via REST/GraphQL/WebSocket APIs and custom SDKs (Python, TypeScript) Act as trusted technical advisor to customers, shaping AI strategy and guiding roadmap decisions from concept to production Data Strategy & MLOps Infrastructure Design data architectures with efficient processing pipelines and ingestion workflows for training and inference on aion's platform Implement RAG systems with vector databases optimizing embedding strategies, chunk sizes, and retrieval methods Prepare and validate datasets for fine-tuning, evaluation, and synthetic data generation Work with other MLEs, MLOps, SREs to carry out model deployment and productionization Observability, Evaluation & Production Operations Implement LLM and agents observability and monitoring tracking token usage, latency, costs, and quality metrics across deployments on aion's infrastructure Instrument applications to trace LLM calls, retrieval operations, agent actions, and data flows Build evaluation frameworks with offline benchmarks (accuracy, relevance, safety metrics) and online monitoring (user feedback, drift detection) Technical Skills & Experience 6-8+ years of hands-on experience building production AI/ML systems, with 3-4+ years deploying LLM applications to production Multimodal AI expertise practical experience building voice agents, vision systems, or conversational AI serving real users Strong LLM foundations hands-on with modern foundation models including fine-tuning, prompt engineering, and evaluation methodologies Agent framework proficiency production experience with LangChain, LlamaIndex, or similar orchestration frameworks Voice AI platform experience built real-time conversational systems with production STT/TTS integration Proficiency in Python (production-grade, async programming, type hints) and JavaScript/TypeScript (full-stack development) RAG implementation experience built retrieval-augmented generation systems with vector databases MLOps & deployment hands-on with Docker, Kubernetes, CI/CD pipelines, and infrastructure-as-code Cloud platforms experience with AWS, Azure, or GCP for ML workloads and infrastructure management Exceptional communication ability to explain complex AI concepts clearly to both technical and business stakeholders Customer-facing experience in Solutions Architecture, Technical Account Management, or Pre-Sales Engineering is highly desirable Computer vision experience working with video processing, object detection, or vision-language models is a plus Model fine-tuning practical experience with LoRA/QLoRA, supervised fine-tuning, or RLHF workflows is a plus Inference optimization experience with vLLM, TensorRT-LLM, Triton, or model quantization techniques is desirable Observability tooling practical experience with LLM monitoring, tracing, and evaluation frameworks is a strong plus Familiarity with WebRTC, real-time streaming protocols, and low-latency media processing Preferred Attributes: Founder-level ownership and bias for action. Strong strategic thinking and ability to connect technical decisions to business impact. Excellent communication and mentoring skills. Thrives in ambiguity, fast-paced environments, and early-stage startup culture. Why Join aion? Work directly with high-pedigree founders shaping technical and product strategy. Build infrastructure powering the future of AI compute globally. Significant ownership and impact with equity reflective of your contributions. Competitive compensation, flexible work options, and wellness benefits
17/07/2026
Full time
About aion Aion is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, Aion takes them from concept to production on a single, unified platform. We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters. We're a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we're building the next generation of enterprise AI and we're looking for exceptional people to help us scale. Who You Are You're a hands-on AI engineer with 3-5+ years of experience building production-grade multimodal AI systems and LLM applications. Your responsibilities mirror those of a hands-on AI startup CTO you work in small teams to own delivery of high-stakes customer projects, embedding directly at client sites to architect, build, and deploy intelligent agent solutions. You're equally comfortable writing production code, presenting technical solutions to C-level executives, and debugging complex AI systems on factory floors or in customer data centers. You've shipped voice agents, video processing systems, or conversational AI to production. You thrive translating ambiguous business requirements into concrete technical solutions that create measurable impact. You're comfortable working across the full AI deployment lifecycle from use case discovery and solution architecture to multimodal agent development, MLOps pipeline implementation, and production optimization. You understand what makes agents perform well in production and how to systematically improve quality through observability and evaluation. Experience with voice AI platforms, RAG systems, and LLM orchestration frameworks is highly desirable. You bring exceptional communication skills, customer empathy, and the drive to build AI solutions that transform enterprise operations globally. What You'll Do Customer Engagement & Multimodal Agent Development Work directly at customer sites from factory floors to executive offices conducting discovery workshops and technical assessments to identify high-impact AI opportunities Design and architect end-to-end multimodal agent systems (voice + video + text) that leverage aion's distributed GPU infrastructure and managed services Build production-grade voice AI systems using STT, TTS APIs, and LLMs deployed on aion's platform Develop vision-enabled agents processing real-time video streams using computer vision pipelines on aion's infrastructure Implement sophisticated multi-agent orchestration with(or similar) frameworks like LangChain or LlamaIndex-enabling tool use, memory management, and autonomous task completion Rapidly prototype POCs in 2-4 weeks, coding alongside client teams to validate concepts and iterate based on feedback Optimize for sub-500ms latency, natural conversation flow, turn detection, and interruption handling in real-time systems Integrate agents directly into customer codebases via REST/GraphQL/WebSocket APIs and custom SDKs (Python, TypeScript) Act as trusted technical advisor to customers, shaping AI strategy and guiding roadmap decisions from concept to production Data Strategy & MLOps Infrastructure Design data architectures with efficient processing pipelines and ingestion workflows for training and inference on aion's platform Implement RAG systems with vector databases optimizing embedding strategies, chunk sizes, and retrieval methods Prepare and validate datasets for fine-tuning, evaluation, and synthetic data generation Work with other MLEs, MLOps, SREs to carry out model deployment and productionization Observability, Evaluation & Production Operations Implement LLM and agents observability and monitoring tracking token usage, latency, costs, and quality metrics across deployments on aion's infrastructure Instrument applications to trace LLM calls, retrieval operations, agent actions, and data flows Build evaluation frameworks with offline benchmarks (accuracy, relevance, safety metrics) and online monitoring (user feedback, drift detection) Technical Skills & Experience 6-8+ years of hands-on experience building production AI/ML systems, with 3-4+ years deploying LLM applications to production Multimodal AI expertise practical experience building voice agents, vision systems, or conversational AI serving real users Strong LLM foundations hands-on with modern foundation models including fine-tuning, prompt engineering, and evaluation methodologies Agent framework proficiency production experience with LangChain, LlamaIndex, or similar orchestration frameworks Voice AI platform experience built real-time conversational systems with production STT/TTS integration Proficiency in Python (production-grade, async programming, type hints) and JavaScript/TypeScript (full-stack development) RAG implementation experience built retrieval-augmented generation systems with vector databases MLOps & deployment hands-on with Docker, Kubernetes, CI/CD pipelines, and infrastructure-as-code Cloud platforms experience with AWS, Azure, or GCP for ML workloads and infrastructure management Exceptional communication ability to explain complex AI concepts clearly to both technical and business stakeholders Customer-facing experience in Solutions Architecture, Technical Account Management, or Pre-Sales Engineering is highly desirable Computer vision experience working with video processing, object detection, or vision-language models is a plus Model fine-tuning practical experience with LoRA/QLoRA, supervised fine-tuning, or RLHF workflows is a plus Inference optimization experience with vLLM, TensorRT-LLM, Triton, or model quantization techniques is desirable Observability tooling practical experience with LLM monitoring, tracing, and evaluation frameworks is a strong plus Familiarity with WebRTC, real-time streaming protocols, and low-latency media processing Preferred Attributes: Founder-level ownership and bias for action. Strong strategic thinking and ability to connect technical decisions to business impact. Excellent communication and mentoring skills. Thrives in ambiguity, fast-paced environments, and early-stage startup culture. Why Join aion? Work directly with high-pedigree founders shaping technical and product strategy. Build infrastructure powering the future of AI compute globally. Significant ownership and impact with equity reflective of your contributions. Competitive compensation, flexible work options, and wellness benefits
Tripledot Studios is one of the largest independent mobile games companies in the world. We are a multi-award-winning organisation, with a global 2,500+ strong team across 12 studios. Our expanded portfolio includes some of the biggest titles in mobile gaming, collectively reaching top chart positions around the world and engaging over 25 million daily active users. Tripledot's guiding principle is that when people love what they do, what they do will be loved by others. We're building a company we're proud of. One filled with driven, incredibly smart and detail-orientated people, who LOVE making games. Our ambition is to be the most successful games company in the world, and we're just getting started. Group role reporting to the IT team The role will focus on AI/ML infrastructure for various group use cases from proprietary models training to efficient 3rd party tool running environments etc. This role will be working with all Group central teams (Marketing, Finance, HR, AI, IT). Role Overview You will build and maintain the high-performance infrastructure required to train and deploy AI models that impact millions of players in real-time, as well as improve productivity of everyone at Tripledot. You will serve as a bridge between the technical solutions that are created by our teams and our live game engines, as well as creating and maintaining infrastructure for internal IT needs. Within the group AI functions you'll be working with other AI / ML engineers, data engineers, analysts and product owners. Within the various studios and other central teams, you will interact with data, engineering, and product teams. Within group IT, you'll partner with Engineering, TechOps, and Security to deliver the infrastructure and tooling that powers our central business functions: gaming, finance, marketing, legal, people ops, and beyond. The first initiative you'll be taking part of is the expansion of a data/ML Platform to support ML engineers and data scientists to easily deploy their solutions, and enabling delivery of key projects like LTV and Ads Optimization. Key Responsibilities Improve and maintain a scalable, speedy and reliable data and ML platform to support AI/ML initiatives within group AI, ensuring models move seamlessly from research to production. Support group IT to provide reliable access to open source AI models and ensure safe reliable access to AI productivity tools. Create and maintain proper monitoring and alerting tools to ensure our systems can provide the correct SLA and SLOs defined by the stakeholders. Implement and advocate for engineering best practices, including CI/CD, infrastructure as code like Terraform, usage of version control, testing, observability, while keeping costs in mind. Ensure standardized cross-studio access & security to enable timely data access and ingestion (AWS and Google Cloud). Enable the teams with different environments for testing new setups, tools, without disrupting the day-to-day operations of the team and production workflows. Track usage for all our deployed applications, and identify areas of improvement, making the best use of resources. Keep up with the relevant technologies, best practices, especially related to AI productivity tools, continuously emerging in the industry. What we are looking for 5+ years in the industry as a DevOps, SRE Engineer or or Platform Engineer, ideally in gaming, mobile apps, or other high-scale digital products. Strong hands on experience with Kubernetes in production - not just running workloads on it, but operating it. Cost aware infrastructure decision making. Solid Terraform (or OpenTofu) experience, with a track record of keeping IaC sustainable as it grows. Proven experience in delivering data and AI/ML solutions in production for both AWS and a working knowledge of GCP or willingness to come to speed quickly. Bonus if this experience is within the gaming industry. Comfortable owning CI/CD pipelines with common tools (GitHub Actions, GitLab CI, ArgoCD, Jenkins, or similar). Hands on experience with cloud and Kubernetes security fundamentals, IAM/RBAC, secrets management (ex. Vault, AWS Secrets Manager, External Secrets), network policies, and integrating security checks into CI/CD pipelines. Strong instincts for observability, monitoring, and alerting, you've built dashboards and alerts that teams actually rely on, and you know the difference between a useful page and noise. Hands on with tools like Prometheus, Grafana, Datadog, CloudWatch, or similar. Solid incident response experience. The current data and AI/ML stack uses open source tools like Airflow, Trino, Spark, and Kubeflow. Familiarity with deploying these tools, as well as tweaking them for improved performance, is a bonus. Understanding of ML Ops best practices and common architectures is also a bonus. Hands on knowledge of Python and/or other scripting languages. Experience creating infrastructure for both traditional and modern agentic data intensive systems is a bonus. Focus on innovation, coupled with a mindset of continuous learning and curiosity to explore emerging AI technologies. The successful candidate will have an agile, hands on approach to prototyping and validation, and ability to Get Stuff Done in a fast paced environment. Excellent communication and collaboration skills necessary for working effectively with both technical and non technical teams. Understanding how to drive results with key business stakeholders. You will be part of a fun mobile gaming company aiming to embrace the future of AI driven creativity and exploring where the industry is moving. You will be instrumental in shaping the backbone of the AI/ML and IT systems that will power solutions that will spread throughout the whole group. You will operate in an environment that values an experimental mindset, focusing on learning opportunities and pioneering generative game creation. Working at Tripledot (in London) 25 days holiday: Enjoy 25 days of paid holiday, in addition to bank holidays to relax and refresh throughout the year. Hybrid Working: We work in the office 3 days a week, Tuesdays and Wednesdays, and a third day of your choice. 20 days fully remote working: Work from anywhere in the world, 20 days of the year. Regular company events and rewards: Join in regular events and rewards that celebrate cultural events, our achievements and our team spirit. Recent highlights have been: Thames River Cruise, London Dungeon and Summer Parties in Regents Park. Continuous Professional Development: Propel your career with continuous opportunities for professional development. Private Medical Cover: Have peace of mind with private medical cover, ensuring your health is in good hands. Life Assurance & Critical Illness Cover: Financial protection for you and your loved ones. Health Cash Plan: Benefit from a health cash plan that contributes to your medical expenses. Dental Cover: Flash your best smile with our dental cover. Family Forming Support: Receive vital support on your family forming/ fertility journey with our support program subject to policy Employee Assistance Program: Anytime you need it, tap into confidential, caring support with our Employee Assistance Program, always here to lend an ear and a helping hand. Cycle to Work Scheme: Make the most of our Cycle to Work Scheme for a green and healthy commute. Pension Plan: Secure your future with our contributory pension plan.
16/07/2026
Full time
Tripledot Studios is one of the largest independent mobile games companies in the world. We are a multi-award-winning organisation, with a global 2,500+ strong team across 12 studios. Our expanded portfolio includes some of the biggest titles in mobile gaming, collectively reaching top chart positions around the world and engaging over 25 million daily active users. Tripledot's guiding principle is that when people love what they do, what they do will be loved by others. We're building a company we're proud of. One filled with driven, incredibly smart and detail-orientated people, who LOVE making games. Our ambition is to be the most successful games company in the world, and we're just getting started. Group role reporting to the IT team The role will focus on AI/ML infrastructure for various group use cases from proprietary models training to efficient 3rd party tool running environments etc. This role will be working with all Group central teams (Marketing, Finance, HR, AI, IT). Role Overview You will build and maintain the high-performance infrastructure required to train and deploy AI models that impact millions of players in real-time, as well as improve productivity of everyone at Tripledot. You will serve as a bridge between the technical solutions that are created by our teams and our live game engines, as well as creating and maintaining infrastructure for internal IT needs. Within the group AI functions you'll be working with other AI / ML engineers, data engineers, analysts and product owners. Within the various studios and other central teams, you will interact with data, engineering, and product teams. Within group IT, you'll partner with Engineering, TechOps, and Security to deliver the infrastructure and tooling that powers our central business functions: gaming, finance, marketing, legal, people ops, and beyond. The first initiative you'll be taking part of is the expansion of a data/ML Platform to support ML engineers and data scientists to easily deploy their solutions, and enabling delivery of key projects like LTV and Ads Optimization. Key Responsibilities Improve and maintain a scalable, speedy and reliable data and ML platform to support AI/ML initiatives within group AI, ensuring models move seamlessly from research to production. Support group IT to provide reliable access to open source AI models and ensure safe reliable access to AI productivity tools. Create and maintain proper monitoring and alerting tools to ensure our systems can provide the correct SLA and SLOs defined by the stakeholders. Implement and advocate for engineering best practices, including CI/CD, infrastructure as code like Terraform, usage of version control, testing, observability, while keeping costs in mind. Ensure standardized cross-studio access & security to enable timely data access and ingestion (AWS and Google Cloud). Enable the teams with different environments for testing new setups, tools, without disrupting the day-to-day operations of the team and production workflows. Track usage for all our deployed applications, and identify areas of improvement, making the best use of resources. Keep up with the relevant technologies, best practices, especially related to AI productivity tools, continuously emerging in the industry. What we are looking for 5+ years in the industry as a DevOps, SRE Engineer or or Platform Engineer, ideally in gaming, mobile apps, or other high-scale digital products. Strong hands on experience with Kubernetes in production - not just running workloads on it, but operating it. Cost aware infrastructure decision making. Solid Terraform (or OpenTofu) experience, with a track record of keeping IaC sustainable as it grows. Proven experience in delivering data and AI/ML solutions in production for both AWS and a working knowledge of GCP or willingness to come to speed quickly. Bonus if this experience is within the gaming industry. Comfortable owning CI/CD pipelines with common tools (GitHub Actions, GitLab CI, ArgoCD, Jenkins, or similar). Hands on experience with cloud and Kubernetes security fundamentals, IAM/RBAC, secrets management (ex. Vault, AWS Secrets Manager, External Secrets), network policies, and integrating security checks into CI/CD pipelines. Strong instincts for observability, monitoring, and alerting, you've built dashboards and alerts that teams actually rely on, and you know the difference between a useful page and noise. Hands on with tools like Prometheus, Grafana, Datadog, CloudWatch, or similar. Solid incident response experience. The current data and AI/ML stack uses open source tools like Airflow, Trino, Spark, and Kubeflow. Familiarity with deploying these tools, as well as tweaking them for improved performance, is a bonus. Understanding of ML Ops best practices and common architectures is also a bonus. Hands on knowledge of Python and/or other scripting languages. Experience creating infrastructure for both traditional and modern agentic data intensive systems is a bonus. Focus on innovation, coupled with a mindset of continuous learning and curiosity to explore emerging AI technologies. The successful candidate will have an agile, hands on approach to prototyping and validation, and ability to Get Stuff Done in a fast paced environment. Excellent communication and collaboration skills necessary for working effectively with both technical and non technical teams. Understanding how to drive results with key business stakeholders. You will be part of a fun mobile gaming company aiming to embrace the future of AI driven creativity and exploring where the industry is moving. You will be instrumental in shaping the backbone of the AI/ML and IT systems that will power solutions that will spread throughout the whole group. You will operate in an environment that values an experimental mindset, focusing on learning opportunities and pioneering generative game creation. Working at Tripledot (in London) 25 days holiday: Enjoy 25 days of paid holiday, in addition to bank holidays to relax and refresh throughout the year. Hybrid Working: We work in the office 3 days a week, Tuesdays and Wednesdays, and a third day of your choice. 20 days fully remote working: Work from anywhere in the world, 20 days of the year. Regular company events and rewards: Join in regular events and rewards that celebrate cultural events, our achievements and our team spirit. Recent highlights have been: Thames River Cruise, London Dungeon and Summer Parties in Regents Park. Continuous Professional Development: Propel your career with continuous opportunities for professional development. Private Medical Cover: Have peace of mind with private medical cover, ensuring your health is in good hands. Life Assurance & Critical Illness Cover: Financial protection for you and your loved ones. Health Cash Plan: Benefit from a health cash plan that contributes to your medical expenses. Dental Cover: Flash your best smile with our dental cover. Family Forming Support: Receive vital support on your family forming/ fertility journey with our support program subject to policy Employee Assistance Program: Anytime you need it, tap into confidential, caring support with our Employee Assistance Program, always here to lend an ear and a helping hand. Cycle to Work Scheme: Make the most of our Cycle to Work Scheme for a green and healthy commute. Pension Plan: Secure your future with our contributory pension plan.