it job board logo
  • Home
  • Find IT Jobs
  • Register CV
  • Career Advice
  • Contact us
  • Employers
    • Register as Employer
    • Pricing Plans
  • Recruiting? Post a job
  • Sign in
  • Sign up
  • Home
  • Find IT Jobs
  • Register CV
  • Career Advice
  • Contact us
  • Employers
    • Register as Employer
    • Pricing Plans
Sorry, that job is no longer available. Here are some results that may be similar to the job you were looking for.

20 jobs found

Email me jobs like this
Refine Search
Current Search
distributed scheduling workload orchestration engineer
Platform Engineer
The Fidelis Partnership
The Fidelis Partnership is a leading privately-owned, Bermuda-based Managing General Underwriter, which, through its subsidiaries, is a global underwriter of property, bespoke and specialty insurance and reinsurance products. The Fidelis Partnership is one of the largest Managing General Underwriters globally and its operations also include outwards reinsurance, claims handling, exposure management and portfolio analytics. The Fidelis Partnership also sponsors and incubates specialist MGAs through its Pine Walk platform. The Fidelis Partnership is separately owned and managed from the ownership and management of Fidelis Insurance Group. Across product lines and geographies, we focus on three diversified pillars: reinsurance, specialty and bespoke solutions. We are truly diversified. Our long-standing partnerships with capital providers and quota share partners make us nimble. Our breadth of expertise and capabilities deliver outstanding market returns. The role The Analytics Product Engineering team at The Fidelis Partnership (TFP) builds and manages a bespoke analytical platform that powers our partnership-driven business model. We combine actuarial expertise with advanced technology to deliver innovative, scalable solutions. The Platform Engineer partners with the Analytics product engineering team to design, build and operate platform capabilities that enable our market leading high-performance, distributed computing products to run reliably, securely and efficiently at scale. This includes infrastructure automation, runtime orchestration, CI/CD enablement, observability and performance optimisation. You will work closely with software engineers, data scientists, product owners, delivery leads, and central IT to define and deliver the platform roadmap, balancing speed, control and resilience in a regulated environment. Key accountabilities Define and execute the platform engineering strategy and roadmap aligned to analytics product needs. Design and operate distributed runtime environments including compute orchestration and workload scheduling. Implement Infrastructure as Code (IaC) and automated environment provisioning. Build and maintain CI/CD pipelines supporting distributed systems and shared components. Implement observability through logging, metrics and alerting, improving platform reliability and debuggability. Monitor and optimise system performance, throughput and resource utilisation. Ensure platform security, access control and compliance with internal standards. Own and maintain operational documentation, runbooks and support procedures to ensure continuity of service. Establish knowledge-sharing and cross-training practices to reduce single points of failure. Ensure key platform processes are documented, repeatable and transferable across the team. Support resilience through shared ownership of critical platform components, releases and incident response. Collaborate with engineering squads and stakeholders to drive adoption of platform capabilities and standards. Skills & experience At least 7 years of experience in platform engineering, DevOps or Site Reliability Engineering supporting distributed, high performance compute systems, ideally with Microsoft HPC in a Windows Server environment. Strong experience with hybrid, bare metal and virtualized infrastructure environments. Strong understanding of networking, security and access control principles and components, including Microsoft Entra (Active Directory). Extensive applied experience of observability practices including logging, monitoring and alerting (e.g. SolarWinds). Practical experience of commissioning and maintaining Microsoft SQL Server installations. Experience designing, deploying and maintaining containerisation and orchestration solutions (e.g. Docker, Kubernetes) a significant plus. Experience of using Infrastructure as Code tools such as Terraform or Bicep for definition and management of environments a significant plus. Solid experience of provisioning and managing storage, backup/restore and Disaster Recovery solutions. Strong communication and stakeholder engagement skills. Diversity, Equity & Inclusion Insofar as possible, we aim to ensure the composition of our workforce reflects the make-up of the local community We have specific programmes in all our offices to support diversity within the hiring process, e.g. internship and scholarship award programmes This is a particular focus in Bermuda, where we engage actively with local organisations to source diverse talent and provide coaching/mentoring for underrepresented groups We aim to maintain a focus on equal opportunities across all stages of hiring process We measure and minimise the pay gap where possible. To ensure that all candidates have a fair opportunity to show their abilities during the recruitment process, adjustments may be required. If your physical or mental health or disability may necessitate an adjustment, please contact to discuss. All information relating to your health or disability will be treated in accordance with our data protection policy. The Leader in Bespoke & Specialty insurance
25/07/2026
Full time
The Fidelis Partnership is a leading privately-owned, Bermuda-based Managing General Underwriter, which, through its subsidiaries, is a global underwriter of property, bespoke and specialty insurance and reinsurance products. The Fidelis Partnership is one of the largest Managing General Underwriters globally and its operations also include outwards reinsurance, claims handling, exposure management and portfolio analytics. The Fidelis Partnership also sponsors and incubates specialist MGAs through its Pine Walk platform. The Fidelis Partnership is separately owned and managed from the ownership and management of Fidelis Insurance Group. Across product lines and geographies, we focus on three diversified pillars: reinsurance, specialty and bespoke solutions. We are truly diversified. Our long-standing partnerships with capital providers and quota share partners make us nimble. Our breadth of expertise and capabilities deliver outstanding market returns. The role The Analytics Product Engineering team at The Fidelis Partnership (TFP) builds and manages a bespoke analytical platform that powers our partnership-driven business model. We combine actuarial expertise with advanced technology to deliver innovative, scalable solutions. The Platform Engineer partners with the Analytics product engineering team to design, build and operate platform capabilities that enable our market leading high-performance, distributed computing products to run reliably, securely and efficiently at scale. This includes infrastructure automation, runtime orchestration, CI/CD enablement, observability and performance optimisation. You will work closely with software engineers, data scientists, product owners, delivery leads, and central IT to define and deliver the platform roadmap, balancing speed, control and resilience in a regulated environment. Key accountabilities Define and execute the platform engineering strategy and roadmap aligned to analytics product needs. Design and operate distributed runtime environments including compute orchestration and workload scheduling. Implement Infrastructure as Code (IaC) and automated environment provisioning. Build and maintain CI/CD pipelines supporting distributed systems and shared components. Implement observability through logging, metrics and alerting, improving platform reliability and debuggability. Monitor and optimise system performance, throughput and resource utilisation. Ensure platform security, access control and compliance with internal standards. Own and maintain operational documentation, runbooks and support procedures to ensure continuity of service. Establish knowledge-sharing and cross-training practices to reduce single points of failure. Ensure key platform processes are documented, repeatable and transferable across the team. Support resilience through shared ownership of critical platform components, releases and incident response. Collaborate with engineering squads and stakeholders to drive adoption of platform capabilities and standards. Skills & experience At least 7 years of experience in platform engineering, DevOps or Site Reliability Engineering supporting distributed, high performance compute systems, ideally with Microsoft HPC in a Windows Server environment. Strong experience with hybrid, bare metal and virtualized infrastructure environments. Strong understanding of networking, security and access control principles and components, including Microsoft Entra (Active Directory). Extensive applied experience of observability practices including logging, monitoring and alerting (e.g. SolarWinds). Practical experience of commissioning and maintaining Microsoft SQL Server installations. Experience designing, deploying and maintaining containerisation and orchestration solutions (e.g. Docker, Kubernetes) a significant plus. Experience of using Infrastructure as Code tools such as Terraform or Bicep for definition and management of environments a significant plus. Solid experience of provisioning and managing storage, backup/restore and Disaster Recovery solutions. Strong communication and stakeholder engagement skills. Diversity, Equity & Inclusion Insofar as possible, we aim to ensure the composition of our workforce reflects the make-up of the local community We have specific programmes in all our offices to support diversity within the hiring process, e.g. internship and scholarship award programmes This is a particular focus in Bermuda, where we engage actively with local organisations to source diverse talent and provide coaching/mentoring for underrepresented groups We aim to maintain a focus on equal opportunities across all stages of hiring process We measure and minimise the pay gap where possible. To ensure that all candidates have a fair opportunity to show their abilities during the recruitment process, adjustments may be required. If your physical or mental health or disability may necessitate an adjustment, please contact to discuss. All information relating to your health or disability will be treated in accordance with our data protection policy. The Leader in Bespoke & Specialty insurance
Data Engineer (SC Cleared)
scrumconnect ltd City, Newcastle Upon Tyne
Apache Spark Python AWS Cloud Data Pipelines A hands-on data engineering role within a large-scale cloud data programme, responsible for building, maintaining, and troubleshooting data pipelines using Apache Spark, PySpark, Apache Airflow, and a broad suite of AWS services. You will apply strong analytical and engineering skills to deliver trusted, well-governed data assets in a modern, cloud-native environment. About Scrumconnect Scrumconnect is a leading UK technology consultancy delivering digital transformation across public and private sectors, contributing to over 20% of the UK's major citizen-facing public services. We specialise in cloud engineering, data platforms, and agile delivery, helping clients build scalable, secure, and user-centred digital solutions that create real impact. Active SC clearance is a mandatory, non-negotiable requirement. Candidates must hold current, in-date Security Check (SC) clearance at the time of application. Sponsorship is not available. Applications without active SC clearance will not be considered. Working arrangement: This role is hybrid. Candidates must be willing and able to travel to the Newcastle office three days per week. Remaining days may be worked remotely from anywhere in the UK. About the role You will work as a Data Engineer on a complex, cloud-based data programme - designing, building, and maintaining data pipelines that process large volumes of data across a modern AWS-native stack. Using Apache Spark and PySpark for distributed data processing, Apache Airflow for orchestration, and a range of AWS services for storage, compute, and analytics, you will help deliver reliable, well-governed data assets to downstream users. You will apply strong data analysis skills to identify root causes of data issues, work with dimensional data models and slowly changing dimensions, and implement infrastructure as code using Terraform. Familiarity with engineering best practices and the ability to translate customer expectations into applied technical functionality are key to success in this role. Key responsibilitiesData pipeline development Build and maintain scalable data pipelines using Apache Spark and PySpark, processing and transforming large datasets across distributed cloud infrastructure. Workflow orchestration Configure and manage Apache Airflow DAGs for task orchestration, ensuring reliable scheduling, monitoring, and execution of data processing workflows. Root cause analysis Perform data analysis to identify and resolve root causes of pipeline failures and data quality issues - including reviewing EMR output logs and CloudWatch metrics. Data modelling Apply understanding of dimensional data models and slowly changing dimensions (SCD) to design and maintain well-structured, analytically trusted data assets. Infrastructure as code Provision and manage cloud infrastructure using Terraform. Containerise solutions using Docker and manage deployments through GitLab CI/CD pipelines and release tagging. Security & encryption Apply understanding of both Server Side and client-side encryption patterns within AWS. Work within IAM policies and data governance standards appropriate to a regulated government environment. Technical skills requiredLanguages & analytics Python - primary language for pipeline development and data processing SQL - used for querying, transformation, and validation across data stores PySpark - for distributed data processing using Apache Spark on AWS EMR Familiarity with basic data structures for constructing robust, scalable solutions Data processing & orchestration Apache Spark - understanding of distributed data processing architecture and execution Apache Airflow - configuring DAGs and managing task orchestration at scale Jupyter Notebooks - for exploratory data analysis and pipeline prototyping Understanding of dimensional data models and slowly changing dimensions (SCD Types 1, 2, 3) Data analysis skills to identify root cause of issues within pipelines and data assets AWS services Amazon EMR - running Spark workloads and reviewing output logs Amazon Athena - ad hoc querying of data in S3 Amazon Textract and Comprehend - familiarity with AI/ML document extraction and NLP services AWS S3, IAM, CloudWatch, EC2, ECR - core platform services used day-to-day AWS console proficiency - navigating, configuring, and monitoring services Understanding of Server Side and client-side encryption within AWS Infrastructure, DevOps & delivery Terraform - Infrastructure as Code for provisioning and managing AWS environments Docker - containerisation of data engineering solutions GitLab - source code management, CI/CD pipeline configuration, release tagging, and component versioning Familiarity with engineering best practices Ability to translate customer expectations into applied, functional technical solutions Technology stack at a glance PythonPySparkSQLApache SparkApache AirflowJupyter NotebooksDimensional modelling/SCDAWS EMRAmazon AthenaAWS S3AWS IAMAWS CloudWatchAWS EC2/ECRAmazon TextractAmazon ComprehendTerraformDockerGitLab CI/CDGitLab Tags
24/07/2026
Apache Spark Python AWS Cloud Data Pipelines A hands-on data engineering role within a large-scale cloud data programme, responsible for building, maintaining, and troubleshooting data pipelines using Apache Spark, PySpark, Apache Airflow, and a broad suite of AWS services. You will apply strong analytical and engineering skills to deliver trusted, well-governed data assets in a modern, cloud-native environment. About Scrumconnect Scrumconnect is a leading UK technology consultancy delivering digital transformation across public and private sectors, contributing to over 20% of the UK's major citizen-facing public services. We specialise in cloud engineering, data platforms, and agile delivery, helping clients build scalable, secure, and user-centred digital solutions that create real impact. Active SC clearance is a mandatory, non-negotiable requirement. Candidates must hold current, in-date Security Check (SC) clearance at the time of application. Sponsorship is not available. Applications without active SC clearance will not be considered. Working arrangement: This role is hybrid. Candidates must be willing and able to travel to the Newcastle office three days per week. Remaining days may be worked remotely from anywhere in the UK. About the role You will work as a Data Engineer on a complex, cloud-based data programme - designing, building, and maintaining data pipelines that process large volumes of data across a modern AWS-native stack. Using Apache Spark and PySpark for distributed data processing, Apache Airflow for orchestration, and a range of AWS services for storage, compute, and analytics, you will help deliver reliable, well-governed data assets to downstream users. You will apply strong data analysis skills to identify root causes of data issues, work with dimensional data models and slowly changing dimensions, and implement infrastructure as code using Terraform. Familiarity with engineering best practices and the ability to translate customer expectations into applied technical functionality are key to success in this role. Key responsibilitiesData pipeline development Build and maintain scalable data pipelines using Apache Spark and PySpark, processing and transforming large datasets across distributed cloud infrastructure. Workflow orchestration Configure and manage Apache Airflow DAGs for task orchestration, ensuring reliable scheduling, monitoring, and execution of data processing workflows. Root cause analysis Perform data analysis to identify and resolve root causes of pipeline failures and data quality issues - including reviewing EMR output logs and CloudWatch metrics. Data modelling Apply understanding of dimensional data models and slowly changing dimensions (SCD) to design and maintain well-structured, analytically trusted data assets. Infrastructure as code Provision and manage cloud infrastructure using Terraform. Containerise solutions using Docker and manage deployments through GitLab CI/CD pipelines and release tagging. Security & encryption Apply understanding of both Server Side and client-side encryption patterns within AWS. Work within IAM policies and data governance standards appropriate to a regulated government environment. Technical skills requiredLanguages & analytics Python - primary language for pipeline development and data processing SQL - used for querying, transformation, and validation across data stores PySpark - for distributed data processing using Apache Spark on AWS EMR Familiarity with basic data structures for constructing robust, scalable solutions Data processing & orchestration Apache Spark - understanding of distributed data processing architecture and execution Apache Airflow - configuring DAGs and managing task orchestration at scale Jupyter Notebooks - for exploratory data analysis and pipeline prototyping Understanding of dimensional data models and slowly changing dimensions (SCD Types 1, 2, 3) Data analysis skills to identify root cause of issues within pipelines and data assets AWS services Amazon EMR - running Spark workloads and reviewing output logs Amazon Athena - ad hoc querying of data in S3 Amazon Textract and Comprehend - familiarity with AI/ML document extraction and NLP services AWS S3, IAM, CloudWatch, EC2, ECR - core platform services used day-to-day AWS console proficiency - navigating, configuring, and monitoring services Understanding of Server Side and client-side encryption within AWS Infrastructure, DevOps & delivery Terraform - Infrastructure as Code for provisioning and managing AWS environments Docker - containerisation of data engineering solutions GitLab - source code management, CI/CD pipeline configuration, release tagging, and component versioning Familiarity with engineering best practices Ability to translate customer expectations into applied, functional technical solutions Technology stack at a glance PythonPySparkSQLApache SparkApache AirflowJupyter NotebooksDimensional modelling/SCDAWS EMRAmazon AthenaAWS S3AWS IAMAWS CloudWatchAWS EC2/ECRAmazon TextractAmazon ComprehendTerraformDockerGitLab CI/CDGitLab Tags
Founding GPU Engineer
Fuse Energy
The Opportunity Demand for high-performance compute capacity across the markets we operate in significantly outpaces what we can currently build, meaning speed to power and reliability are critical to how we scale. This puts CUDA/GPU performance engineering at the center of how Fuse scales its compute infrastructure. Responsibilities Design, implement, and optimise CUDA kernels for high-throughput, latency-sensitive workloads. Profile and tune GPU performance across compute, memory bandwidth, and interconnect (NVLink/PCIe) bottlenecks. Build tooling to correlate GPU cluster power draw and utilisation with real-time energy pricing and grid signals. Optimise multi-GPU and multi-node scaling using NCCL, MPI, or similar communication libraries. Work with data center infrastructure teams on power capping, dynamic voltage/frequency scaling, and workload scheduling strategies that reduce energy cost and carbon intensity. Collaborate with ML/systems engineers to integrate custom kernels into training/inference pipelines. Benchmark against CPU/GPU baselines and drive continuous performance improvements. Contribute to internal libraries, documentation, and best practices for GPU performance engineering. 4+ years of experience writing production CUDA code, or equivalent strong project/industry experience. Deep understanding of GPU architecture (SMs, warps, memory hierarchy, occupancy). Proficiency in C++ and CUDA; experience with Python for tooling/orchestration. Experience with performance profiling tools (Nsight Systems/Compute). Familiarity with multi-GPU/multi-node scaling (NCCL, MPI, RDMA/InfiniBand). Strong grasp of memory optimisation, kernel fusion, and parallel algorithm design. Comfortable working across the stack from low-level kernels to system-level infrastructure. Nice to Have Experience with Triton, cuDNN, cuBLAS, or custom ML inference/training frameworks. Exposure to data center power/thermal management or demand-response systems. Background in HPC, quantitative finance, or large-scale distributed systems. Familiarity with Kubernetes/Slurm for GPU cluster orchestration. Interest or experience in energy markets, grid systems, or sustainability-focused compute. Competitive salary and an equity sign-on bonus. Biannual bonus scheme. Fully expensed tech to match your needs. Breakfast and dinner allowance for office based employees.
23/07/2026
Full time
The Opportunity Demand for high-performance compute capacity across the markets we operate in significantly outpaces what we can currently build, meaning speed to power and reliability are critical to how we scale. This puts CUDA/GPU performance engineering at the center of how Fuse scales its compute infrastructure. Responsibilities Design, implement, and optimise CUDA kernels for high-throughput, latency-sensitive workloads. Profile and tune GPU performance across compute, memory bandwidth, and interconnect (NVLink/PCIe) bottlenecks. Build tooling to correlate GPU cluster power draw and utilisation with real-time energy pricing and grid signals. Optimise multi-GPU and multi-node scaling using NCCL, MPI, or similar communication libraries. Work with data center infrastructure teams on power capping, dynamic voltage/frequency scaling, and workload scheduling strategies that reduce energy cost and carbon intensity. Collaborate with ML/systems engineers to integrate custom kernels into training/inference pipelines. Benchmark against CPU/GPU baselines and drive continuous performance improvements. Contribute to internal libraries, documentation, and best practices for GPU performance engineering. 4+ years of experience writing production CUDA code, or equivalent strong project/industry experience. Deep understanding of GPU architecture (SMs, warps, memory hierarchy, occupancy). Proficiency in C++ and CUDA; experience with Python for tooling/orchestration. Experience with performance profiling tools (Nsight Systems/Compute). Familiarity with multi-GPU/multi-node scaling (NCCL, MPI, RDMA/InfiniBand). Strong grasp of memory optimisation, kernel fusion, and parallel algorithm design. Comfortable working across the stack from low-level kernels to system-level infrastructure. Nice to Have Experience with Triton, cuDNN, cuBLAS, or custom ML inference/training frameworks. Exposure to data center power/thermal management or demand-response systems. Background in HPC, quantitative finance, or large-scale distributed systems. Familiarity with Kubernetes/Slurm for GPU cluster orchestration. Interest or experience in energy markets, grid systems, or sustainability-focused compute. Competitive salary and an equity sign-on bonus. Biannual bonus scheme. Fully expensed tech to match your needs. Breakfast and dinner allowance for office based employees.
AI Inference Engineer
Fuse Energy
Fuse Energy is a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy fast. We're combining first-principles thinking with cutting-edge technology to build a radically better energy system. We raised $210M from top-tier investors including Multicoin, Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, Box Group and strategic angels like Nico Rosberg, the Co-Founder of Solana and GPs behind Meta, Revolut, Spotify, Uber and more. As data centres become one of the largest and fastest-growing sources of electricity demand, Fuse is expanding into high-performance compute infrastructure that sits at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch and we're looking for the founding engineer to own the latter. We're looking for a Founding AI Inference Engineer to define and build how Fuse serves AI inference workloads at scale, reporting directly to the CTO. Where our CUDA and GPU engineering hires own kernel-level and hardware performance, this role owns the layer above it: how models actually get served, scaled, and delivered against committed performance targets. The Opportunity Fuse is seeing significant demand for data centre capacity across the markets we operate in, primarily for inference. Few companies in the world can pair real power delivery with real compute the way Fuse can, which puts inference serving at the heart of how we turn that advantage into the best offering in the market. That's this role. Responsibilities Define Fuse's inference serving strategy and architecture from first principles. Design and build the serving stack: request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference workloads. Own model-level optimisation strategy for serving - deciding where and how to apply quantisation, distillation, speculative decoding, and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers. Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents). Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans. Act as a direct technical owner of inference performance and reliability. Work closely with the CUDA and GPU engineering teams to ensure custom kernels and hardware performance work are integrated cleanly into the serving layer. Set the standards, tooling, and benchmarks this function will run on as it grows. Qualifications 4+ years of experience building or operating large-scale inference serving systems, or equivalent strong project/industry experience. Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding). Strong systems thinking - able to reason about the full path from incoming request to served response across a large cluster. Comfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system. A track record of making high-stakes architecture calls and owning the outcome. Comfort operating without a playbook - this is a founding role shaping a new function around architecture that's still early-stage, not joining an established one. Nice to Have Experience with Triton or custom ML inference/training frameworks. Experience with autoscaling or capacity planning for large-scale inference workloads. Exposure to multi-tenant serving or SLA-driven infrastructure. Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system. Familiarity with Kubernetes/Slurm for cluster orchestration. Interest or experience in energy markets, grid systems, or sustainability-focused compute. Benefits Competitive salary and an equity sign-on bonus. Biannual bonus scheme. Fully expensed tech to match your needs. Breakfast and dinner allowance for office based employees.
23/07/2026
Full time
Fuse Energy is a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy fast. We're combining first-principles thinking with cutting-edge technology to build a radically better energy system. We raised $210M from top-tier investors including Multicoin, Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, Box Group and strategic angels like Nico Rosberg, the Co-Founder of Solana and GPs behind Meta, Revolut, Spotify, Uber and more. As data centres become one of the largest and fastest-growing sources of electricity demand, Fuse is expanding into high-performance compute infrastructure that sits at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch and we're looking for the founding engineer to own the latter. We're looking for a Founding AI Inference Engineer to define and build how Fuse serves AI inference workloads at scale, reporting directly to the CTO. Where our CUDA and GPU engineering hires own kernel-level and hardware performance, this role owns the layer above it: how models actually get served, scaled, and delivered against committed performance targets. The Opportunity Fuse is seeing significant demand for data centre capacity across the markets we operate in, primarily for inference. Few companies in the world can pair real power delivery with real compute the way Fuse can, which puts inference serving at the heart of how we turn that advantage into the best offering in the market. That's this role. Responsibilities Define Fuse's inference serving strategy and architecture from first principles. Design and build the serving stack: request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference workloads. Own model-level optimisation strategy for serving - deciding where and how to apply quantisation, distillation, speculative decoding, and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers. Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents). Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans. Act as a direct technical owner of inference performance and reliability. Work closely with the CUDA and GPU engineering teams to ensure custom kernels and hardware performance work are integrated cleanly into the serving layer. Set the standards, tooling, and benchmarks this function will run on as it grows. Qualifications 4+ years of experience building or operating large-scale inference serving systems, or equivalent strong project/industry experience. Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding). Strong systems thinking - able to reason about the full path from incoming request to served response across a large cluster. Comfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system. A track record of making high-stakes architecture calls and owning the outcome. Comfort operating without a playbook - this is a founding role shaping a new function around architecture that's still early-stage, not joining an established one. Nice to Have Experience with Triton or custom ML inference/training frameworks. Experience with autoscaling or capacity planning for large-scale inference workloads. Exposure to multi-tenant serving or SLA-driven infrastructure. Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system. Familiarity with Kubernetes/Slurm for cluster orchestration. Interest or experience in energy markets, grid systems, or sustainability-focused compute. Benefits Competitive salary and an equity sign-on bonus. Biannual bonus scheme. Fully expensed tech to match your needs. Breakfast and dinner allowance for office based employees.
Founding GPU Engineer
Fuse Energy, LLC
The Opportunity Demand for high-performance compute capacity across the markets we operate in significantly outpaces what we can currently build, meaning speed to power and reliability are critical to how we scale. This puts CUDA/GPU performance engineering at the center of how Fuse scales its compute infrastructure. Responsibilities Design, implement, and optimise CUDA kernels for high-throughput, latency-sensitive workloads. Profile and tune GPU performance across compute, memory bandwidth, and interconnect (NVLink/PCIe) bottlenecks. Build tooling to correlate GPU cluster power draw and utilisation with real-time energy pricing and grid signals. Optimise multi-GPU and multi-node scaling using NCCL, MPI, or similar communication libraries. Work with data center infrastructure teams on power capping, dynamic voltage/frequency scaling, and workload scheduling strategies that reduce energy cost and carbon intensity. Collaborate with ML/systems engineers to integrate custom kernels into training/inference pipelines. Benchmark against CPU/GPU baselines and drive continuous performance improvements. Contribute to internal libraries, documentation, and best practices for GPU performance engineering. 4+ years of experience writing production CUDA code, or equivalent strong project/industry experience. Deep understanding of GPU architecture (SMs, warps, memory hierarchy, occupancy). Proficiency in C++ and CUDA; experience with Python for tooling/orchestration. Experience with performance profiling tools (Nsight Systems/Compute). Familiarity with multi-GPU/multi-node scaling (NCCL, MPI, RDMA/InfiniBand). Strong grasp of memory optimisation, kernel fusion, and parallel algorithm design. Comfortable working across the stack from low-level kernels to system-level infrastructure. Nice to Have Experience with Triton, cuDNN, cuBLAS, or custom ML inference/training frameworks. Exposure to data center power/thermal management or demand-response systems. Background in HPC, quantitative finance, or large-scale distributed systems. Familiarity with Kubernetes/Slurm for GPU cluster orchestration. Interest or experience in energy markets, grid systems, or sustainability-focused compute. Competitive salary and an equity sign-on bonus. Biannual bonus scheme. Fully expensed tech to match your needs. Breakfast and dinner allowance for office based employees.
22/07/2026
Full time
The Opportunity Demand for high-performance compute capacity across the markets we operate in significantly outpaces what we can currently build, meaning speed to power and reliability are critical to how we scale. This puts CUDA/GPU performance engineering at the center of how Fuse scales its compute infrastructure. Responsibilities Design, implement, and optimise CUDA kernels for high-throughput, latency-sensitive workloads. Profile and tune GPU performance across compute, memory bandwidth, and interconnect (NVLink/PCIe) bottlenecks. Build tooling to correlate GPU cluster power draw and utilisation with real-time energy pricing and grid signals. Optimise multi-GPU and multi-node scaling using NCCL, MPI, or similar communication libraries. Work with data center infrastructure teams on power capping, dynamic voltage/frequency scaling, and workload scheduling strategies that reduce energy cost and carbon intensity. Collaborate with ML/systems engineers to integrate custom kernels into training/inference pipelines. Benchmark against CPU/GPU baselines and drive continuous performance improvements. Contribute to internal libraries, documentation, and best practices for GPU performance engineering. 4+ years of experience writing production CUDA code, or equivalent strong project/industry experience. Deep understanding of GPU architecture (SMs, warps, memory hierarchy, occupancy). Proficiency in C++ and CUDA; experience with Python for tooling/orchestration. Experience with performance profiling tools (Nsight Systems/Compute). Familiarity with multi-GPU/multi-node scaling (NCCL, MPI, RDMA/InfiniBand). Strong grasp of memory optimisation, kernel fusion, and parallel algorithm design. Comfortable working across the stack from low-level kernels to system-level infrastructure. Nice to Have Experience with Triton, cuDNN, cuBLAS, or custom ML inference/training frameworks. Exposure to data center power/thermal management or demand-response systems. Background in HPC, quantitative finance, or large-scale distributed systems. Familiarity with Kubernetes/Slurm for GPU cluster orchestration. Interest or experience in energy markets, grid systems, or sustainability-focused compute. Competitive salary and an equity sign-on bonus. Biannual bonus scheme. Fully expensed tech to match your needs. Breakfast and dinner allowance for office based employees.
MLOps Engineer
Lumai Limited Oxford, Oxfordshire
The Opportunity Lumai is redefining how the world computes. We are an ambitious, venture-backed UK startup pioneering a breakthrough AI accelerator for data centers which uses 3D optical compute. Our radical technology uses light to perform computation at orders of magnitude faster speeds and at far greater scales than ever before, all whilst consuming far less energy than traditional approaches. Lumai is unlocking performance and efficiency gains that could transform the economics of AI and compute infrastructure and reshape how intelligence scales globally. If you are passionate about bringing groundbreaking technology to market, and want to be part of a team pushing the boundaries of what is physically possible, Lumai is where you can make it happen. About Lumai Founded in 2022, Lumai is a University of Oxford spinout using optical processing to accelerate large language models (LLMs) and other transformer-based AI systems. The team combines expertise in optical computing, machine learning, and physics. Lumai has already secured over $15 million in investment from leading deep-tech investors like Constructor Capital, IP Group, PhotonVentures and government grants, and is scaling rapidly to deploy the fastest optical compute currently available globally. The Role We are building custom AI hardware and the full-stack software ecosystem to run it. As our first dedicated MLOps Engineer, you will own the infrastructure that takes models from research to silicon-validated production - designing, building, and operating the pipelines, tooling, and platforms that let our AI and hardware teams move fast without breaking things. This is a high-impact, high-ownership role at the intersection of ML research, compiler stacks, and novel hardware. What You'll Do Design and operate end-to-end ML pipelines: data ingest, training, evaluation, quantisation, and deployment onto custom AI accelerator hardware Build and maintain experiment tracking, model registry, and versioning infrastructure (e.g. MLflow, W&B, or equivalent) tuned to our hardware-in-the-loop workflows Own CI/CD for ML: automated testing of model correctness, numerical accuracy, and on-chip performance after every change to models, compilers, or firmware Develop and maintain tooling for benchmarking model inference on custom silicon, including latency, throughput, power, and utilisation metrics Collaborate closely with ML researchers, compiler engineers, and hardware architects to identify and remove bottlenecks across the model-to-chip workflow Instrument and monitor production inference deployments; design alerting and rollback strategies appropriate to hardware-accelerated serving Manage compute resource scheduling across on-premises accelerator clusters and cloud (GPU/CPU) for training and simulation workloads Drive infrastructure-as-code practices: containerisation, orchestration (Kubernetes/Slurm), and reproducible environment management Contribute to the internal developer platform: self-service tooling, documentation, and runbooks that raise engineering productivity across the company What We're Looking For Must-Have 5+ years of software or infrastructure engineering experience, with at least 2 years in an ML or AI-adjacent role Strong Python skills and familiarity with major ML frameworks (PyTorch or JAX); comfortable reading and modifying model code Hands-on experience building and operating ML pipelines in production: data pipelines, training orchestration, evaluation, and serving Experience with experiment tracking and model lifecycle management tools (MLflow, W&B, DVC, or similar) Solid understanding of containerisation (Docker) and orchestration (Kubernetes or Slurm) for distributed compute workloads Infrastructure-as-code mindset: Terraform, Ansible, or equivalent; CI/CD pipelines (GitHub Actions, Jenkins, or similar) Experience with hardware-accelerated compute (CUDA/GPU workflows, profiling, performance tuning) - even if not on custom silicon Strong debugging and observability skills: distributed tracing, logging, metrics dashboards Ability to work effectively in a fast-moving, ambiguous environment where the hardware and software are both being built simultaneously Strong Preference For Experience with custom or novel accelerator hardware (FPGAs, ASICs, NPUs, or research chips) Familiarity with ML compiler stacks: MLIR, LLVM, TVM, XLA, or vendor-specific compilers (NVCC, TensorRT, etc.) Experience with model optimisation techniques: quantisation (INT8/INT4/FP8), pruning, distillation, or mixed-precision training Background in on-chip performance profiling and roofline analysis Exposure to chip bring-up workflows: running early software stacks on pre-silicon simulation or first-silicon hardware Contributions to open-source ML infrastructure or compiler tooling Experience in a deeptech, semiconductor, or hardware startup environment Compensation & Benefits Highly Competitive Salary: We are not saying our salary is a blank check, but let's just say it won't be a source of your stress Share Option Scheme: We are all in this together! We believe in shared success while we build the Lumai of tomorrow Pension Scheme: Plan for retirement with AVIVA Private Health Insurance: We firmly believe that you come first, and a happy you is a healthy you! Look after yourself and your loved ones with AXA Cycle to Work: Spread the cost of a bike, a bike and accessories or just accessories and save on tax L&D Allowance: Stay at the forefront of your field with a £500 annual development budget Subsidised On-site Lunches: Enjoy on-site healthy meals at half the price, as Lumai covers 50% of the cost Holidays: Enjoy some deserved "me time" with 25 days paid holiday (plus bank holidays) per year Socials: Be part of an inclusive community enjoying occasional all-company off-sites, lunches and socials Interview Process Our process is four stages. An initial conversation with our HR team to understand what you want from the role and what we want to Lumai is an equal opportunity employer. We make hiring decisions on merit, scope-fit, and the strength of the working relationship we expect to build with each hire. Applications welcome from candidates of any background. If you are not sure whether you are a fit, send a note anyway.
19/07/2026
Full time
The Opportunity Lumai is redefining how the world computes. We are an ambitious, venture-backed UK startup pioneering a breakthrough AI accelerator for data centers which uses 3D optical compute. Our radical technology uses light to perform computation at orders of magnitude faster speeds and at far greater scales than ever before, all whilst consuming far less energy than traditional approaches. Lumai is unlocking performance and efficiency gains that could transform the economics of AI and compute infrastructure and reshape how intelligence scales globally. If you are passionate about bringing groundbreaking technology to market, and want to be part of a team pushing the boundaries of what is physically possible, Lumai is where you can make it happen. About Lumai Founded in 2022, Lumai is a University of Oxford spinout using optical processing to accelerate large language models (LLMs) and other transformer-based AI systems. The team combines expertise in optical computing, machine learning, and physics. Lumai has already secured over $15 million in investment from leading deep-tech investors like Constructor Capital, IP Group, PhotonVentures and government grants, and is scaling rapidly to deploy the fastest optical compute currently available globally. The Role We are building custom AI hardware and the full-stack software ecosystem to run it. As our first dedicated MLOps Engineer, you will own the infrastructure that takes models from research to silicon-validated production - designing, building, and operating the pipelines, tooling, and platforms that let our AI and hardware teams move fast without breaking things. This is a high-impact, high-ownership role at the intersection of ML research, compiler stacks, and novel hardware. What You'll Do Design and operate end-to-end ML pipelines: data ingest, training, evaluation, quantisation, and deployment onto custom AI accelerator hardware Build and maintain experiment tracking, model registry, and versioning infrastructure (e.g. MLflow, W&B, or equivalent) tuned to our hardware-in-the-loop workflows Own CI/CD for ML: automated testing of model correctness, numerical accuracy, and on-chip performance after every change to models, compilers, or firmware Develop and maintain tooling for benchmarking model inference on custom silicon, including latency, throughput, power, and utilisation metrics Collaborate closely with ML researchers, compiler engineers, and hardware architects to identify and remove bottlenecks across the model-to-chip workflow Instrument and monitor production inference deployments; design alerting and rollback strategies appropriate to hardware-accelerated serving Manage compute resource scheduling across on-premises accelerator clusters and cloud (GPU/CPU) for training and simulation workloads Drive infrastructure-as-code practices: containerisation, orchestration (Kubernetes/Slurm), and reproducible environment management Contribute to the internal developer platform: self-service tooling, documentation, and runbooks that raise engineering productivity across the company What We're Looking For Must-Have 5+ years of software or infrastructure engineering experience, with at least 2 years in an ML or AI-adjacent role Strong Python skills and familiarity with major ML frameworks (PyTorch or JAX); comfortable reading and modifying model code Hands-on experience building and operating ML pipelines in production: data pipelines, training orchestration, evaluation, and serving Experience with experiment tracking and model lifecycle management tools (MLflow, W&B, DVC, or similar) Solid understanding of containerisation (Docker) and orchestration (Kubernetes or Slurm) for distributed compute workloads Infrastructure-as-code mindset: Terraform, Ansible, or equivalent; CI/CD pipelines (GitHub Actions, Jenkins, or similar) Experience with hardware-accelerated compute (CUDA/GPU workflows, profiling, performance tuning) - even if not on custom silicon Strong debugging and observability skills: distributed tracing, logging, metrics dashboards Ability to work effectively in a fast-moving, ambiguous environment where the hardware and software are both being built simultaneously Strong Preference For Experience with custom or novel accelerator hardware (FPGAs, ASICs, NPUs, or research chips) Familiarity with ML compiler stacks: MLIR, LLVM, TVM, XLA, or vendor-specific compilers (NVCC, TensorRT, etc.) Experience with model optimisation techniques: quantisation (INT8/INT4/FP8), pruning, distillation, or mixed-precision training Background in on-chip performance profiling and roofline analysis Exposure to chip bring-up workflows: running early software stacks on pre-silicon simulation or first-silicon hardware Contributions to open-source ML infrastructure or compiler tooling Experience in a deeptech, semiconductor, or hardware startup environment Compensation & Benefits Highly Competitive Salary: We are not saying our salary is a blank check, but let's just say it won't be a source of your stress Share Option Scheme: We are all in this together! We believe in shared success while we build the Lumai of tomorrow Pension Scheme: Plan for retirement with AVIVA Private Health Insurance: We firmly believe that you come first, and a happy you is a healthy you! Look after yourself and your loved ones with AXA Cycle to Work: Spread the cost of a bike, a bike and accessories or just accessories and save on tax L&D Allowance: Stay at the forefront of your field with a £500 annual development budget Subsidised On-site Lunches: Enjoy on-site healthy meals at half the price, as Lumai covers 50% of the cost Holidays: Enjoy some deserved "me time" with 25 days paid holiday (plus bank holidays) per year Socials: Be part of an inclusive community enjoying occasional all-company off-sites, lunches and socials Interview Process Our process is four stages. An initial conversation with our HR team to understand what you want from the role and what we want to Lumai is an equal opportunity employer. We make hiring decisions on merit, scope-fit, and the strength of the working relationship we expect to build with each hire. Applications welcome from candidates of any background. If you are not sure whether you are a fit, send a note anyway.
Software Engineer, General
AION
About aion Aion is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, Aion takes them from concept to production on a single, unified platform. We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters. We're a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we're building the next generation of enterprise AI and we're looking for exceptional people to help us scale. Who You Are You're a solid engineer with 2-4 years of experience building backend systems and platform infrastructure. You write clean, well-abstracted code with proper design patterns and comprehensive test coverage. You're comfortable working on both the Compute Platform (multi-cloud orchestration, resource management) and Inference Platform (model serving, autoscaling) under the guidance of senior engineers and platform leads. You have strong proficiency in Golang and understand how to build maintainable, production-grade distributed systems. You take pride in code quality, enjoy collaborating on low-level designs, and are eager to learn from experienced engineers while contributing meaningfully to critical infrastructure components. You're product-minded, you understand how your technical decisions impact developers using aion's platform and think about the end-to-end user experience. You're a team player comfortable wearing multiple hats one day you're building product features, the next you're joining customer calls to understand their deployment challenges, and the day after you're helping with UI/UX, customer success, documentation and product ops. What You'll Do Platform Development & Implementation Build and maintain platform services across aion's Compute and Inference platforms, working closely with senior engineers and platform leads Implement features for multi-cloud orchestration, resource scheduling, model deployment pipelines, and autoscaling systems Write well-maintained, production-grade code with proper abstractions, design patterns, and comprehensive test coverage Contribute to low-level design (LLD) including service APIs, database schema design, data models, and component interactions Collaborate with senior engineers on high-level design discussions, providing implementation perspectives and feasibility inputs Backend Systems & Distributed Infrastructure Develop RESTful APIs and gRPC services for platform control planes, resource management, and inference serving Design and implement database schemas for storing platform state, resource metadata, billing data, and observability metrics Work with distributed storage systems, message queues (Kafka, RabbitMQ), and databases (PostgreSQL, Redis) to build reliable platform components Build event-driven architectures for asynchronous processing, job scheduling, and platform automation Implement monitoring, logging, and alerting for platform services to ensure production reliability Code Quality & Engineering Excellence Write comprehensive unit tests, integration tests, and end-to-end tests to ensure code reliability Participate in code reviews, providing constructive feedback and learning from senior engineers' perspectives Refactor existing code to improve maintainability, performance, and scalability Document design decisions, API specifications, and operational runbooks for platform services Debug production issues and contribute to incident response and post-mortems Technical Skills & Experience 2-4 years of experience in backend engineering, platform development, or distributed systems Strong proficiency in Golang you write idiomatic Go code with proper error handling, concurrency patterns, and testing Solid understanding of backend systems fundamentals: RESTful APIs, microservices architecture, and API design principles Hands on experience with databases (PostgreSQL, MySQL) including schema design, query optimization, and transactions Familiarity with storage systems (object storage like S3, block storage, distributed file systems) and their use cases Experience working with message queues (Kafka, RabbitMQ, NATS) and event-driven architectures Understanding of distributed systems concepts: consensus, eventual consistency, fault tolerance, and retry mechanisms Experience with containerization (Docker) and basic Kubernetes concepts Knowledge of testing frameworks and practices (unit tests, integration tests, mocking) Familiarity with Git, CI/CD pipelines, and modern development workflows Exposure to cloud platforms (AWS/GCP/Azure) and their core services is a plus Experience with infrastructure-as-code (Terraform) or observability tools (Prometheus, Grafana) is beneficial Bonus/ Good to Have HPC & Cluster Management: Experience handling large-scale HPC clusters using Kubernetes and Slurm for job scheduling, resource allocation, and workload orchestration Data Engineering: Expertise with data pipelines, ETL systems, and large-scale data processing frameworks Systems-Level Programming: Experience with low-level systems programming such as storage systems, Kubernetes operators, OS-level software development, or daemon services (llm-d, system agents) ML Platform Engineering: Experience productionizing ML pipelines, batch job orchestration, model fine-tuning workflows, and Jupyter notebook orchestration systems Enterprise Deployment: Experience platformizing and packaging software for on-premises deployments or customer VPC installations with emphasis on security, compliance, and operational simplicity Preferred Attributes: High ownership, self driven and biased for action. Strong strategic thinking and ability to connect technical decisions to business impact. Excellent communication and mentoring skills. Thrives in ambiguity, fast-paced environments, and early stage startup culture. Why Join aion? Work directly with high pedigree founders shaping technical and product strategy. Build infrastructure powering the future of AI computers globally. Significant ownership and impact with equity reflective of your contributions. Competitive compensation, flexible work options, and wellness benefits
18/07/2026
Full time
About aion Aion is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, Aion takes them from concept to production on a single, unified platform. We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters. We're a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we're building the next generation of enterprise AI and we're looking for exceptional people to help us scale. Who You Are You're a solid engineer with 2-4 years of experience building backend systems and platform infrastructure. You write clean, well-abstracted code with proper design patterns and comprehensive test coverage. You're comfortable working on both the Compute Platform (multi-cloud orchestration, resource management) and Inference Platform (model serving, autoscaling) under the guidance of senior engineers and platform leads. You have strong proficiency in Golang and understand how to build maintainable, production-grade distributed systems. You take pride in code quality, enjoy collaborating on low-level designs, and are eager to learn from experienced engineers while contributing meaningfully to critical infrastructure components. You're product-minded, you understand how your technical decisions impact developers using aion's platform and think about the end-to-end user experience. You're a team player comfortable wearing multiple hats one day you're building product features, the next you're joining customer calls to understand their deployment challenges, and the day after you're helping with UI/UX, customer success, documentation and product ops. What You'll Do Platform Development & Implementation Build and maintain platform services across aion's Compute and Inference platforms, working closely with senior engineers and platform leads Implement features for multi-cloud orchestration, resource scheduling, model deployment pipelines, and autoscaling systems Write well-maintained, production-grade code with proper abstractions, design patterns, and comprehensive test coverage Contribute to low-level design (LLD) including service APIs, database schema design, data models, and component interactions Collaborate with senior engineers on high-level design discussions, providing implementation perspectives and feasibility inputs Backend Systems & Distributed Infrastructure Develop RESTful APIs and gRPC services for platform control planes, resource management, and inference serving Design and implement database schemas for storing platform state, resource metadata, billing data, and observability metrics Work with distributed storage systems, message queues (Kafka, RabbitMQ), and databases (PostgreSQL, Redis) to build reliable platform components Build event-driven architectures for asynchronous processing, job scheduling, and platform automation Implement monitoring, logging, and alerting for platform services to ensure production reliability Code Quality & Engineering Excellence Write comprehensive unit tests, integration tests, and end-to-end tests to ensure code reliability Participate in code reviews, providing constructive feedback and learning from senior engineers' perspectives Refactor existing code to improve maintainability, performance, and scalability Document design decisions, API specifications, and operational runbooks for platform services Debug production issues and contribute to incident response and post-mortems Technical Skills & Experience 2-4 years of experience in backend engineering, platform development, or distributed systems Strong proficiency in Golang you write idiomatic Go code with proper error handling, concurrency patterns, and testing Solid understanding of backend systems fundamentals: RESTful APIs, microservices architecture, and API design principles Hands on experience with databases (PostgreSQL, MySQL) including schema design, query optimization, and transactions Familiarity with storage systems (object storage like S3, block storage, distributed file systems) and their use cases Experience working with message queues (Kafka, RabbitMQ, NATS) and event-driven architectures Understanding of distributed systems concepts: consensus, eventual consistency, fault tolerance, and retry mechanisms Experience with containerization (Docker) and basic Kubernetes concepts Knowledge of testing frameworks and practices (unit tests, integration tests, mocking) Familiarity with Git, CI/CD pipelines, and modern development workflows Exposure to cloud platforms (AWS/GCP/Azure) and their core services is a plus Experience with infrastructure-as-code (Terraform) or observability tools (Prometheus, Grafana) is beneficial Bonus/ Good to Have HPC & Cluster Management: Experience handling large-scale HPC clusters using Kubernetes and Slurm for job scheduling, resource allocation, and workload orchestration Data Engineering: Expertise with data pipelines, ETL systems, and large-scale data processing frameworks Systems-Level Programming: Experience with low-level systems programming such as storage systems, Kubernetes operators, OS-level software development, or daemon services (llm-d, system agents) ML Platform Engineering: Experience productionizing ML pipelines, batch job orchestration, model fine-tuning workflows, and Jupyter notebook orchestration systems Enterprise Deployment: Experience platformizing and packaging software for on-premises deployments or customer VPC installations with emphasis on security, compliance, and operational simplicity Preferred Attributes: High ownership, self driven and biased for action. Strong strategic thinking and ability to connect technical decisions to business impact. Excellent communication and mentoring skills. Thrives in ambiguity, fast-paced environments, and early stage startup culture. Why Join aion? Work directly with high pedigree founders shaping technical and product strategy. Build infrastructure powering the future of AI computers globally. Significant ownership and impact with equity reflective of your contributions. Competitive compensation, flexible work options, and wellness benefits
Sr./Staff Software Engineer (Data Team)
Thehumanoid
Here at Humanoid, we believe in a future where robots amplify human potential. That's why we've set out on a mission to build the world's most capable, commercially-scalable, and safe humanoid robots. We're bringing that mission to life with HMND 01 Alpha - our rapidly developed humanoid platform now running in real industrial pilots - and we're growing the team to take it even further. About the Role We're hiring a Sr./Staff Software Engineer to join our Data Engineering team based in London. As the Data Engineering Lead, we strive to create the world's leading, commercially scalable, safe, and advanced humanoid robots that seamlessly integrate into daily life and amplify human capacity. We're looking for experienced Engineers to help build our data platform from the ground up. You'll have the opportunity to define its architecture, make key decisions, and shape how we handle and process petabyte-scale data in the near future. What You'll Do Build the Capability Factory - an internal platform designed for everyone from software engineers to non-technical operators, enabling the entire organization to teach HMND robots new skills at scale, from raw data all the way to deployed capabilities. Curate, preprocess, and manage large-scale datasets for humanoid robot training - a corpus of robot telemetry growing toward petabyte scale. Design and operate highly scalable data pipelines and the compute infrastructure that powers them, ensuring reliability and throughput as data volume and team demands grow. Ensure the quality, accuracy, and consistency of training data across multiple concurrent projects and robot platforms. Collaborate with machine learning teams to shape the Capability Factory, streamline MLOps, and build the evaluation workflows that close the loop between training runs and real-world robot performance. Build data warehouse solutions and BI dashboards that give stakeholders across the organization clear visibility into data collection, model progress, and operational health. Establish and uphold best practices for data management - versioning, access control, security, and compliance. What We're Looking For 5+ years of software engineering experience, with a track record of owning and delivering complex systems end-to-end, not just contributing to them. Strong backend engineering - designing and operating production-grade APIs and services: clean data modeling, reliable error handling, performance under load. Data engineering at PB+ scale - building and maintaining pipelines that move, transform, and validate large volumes of data reliably; understanding of batch and streaming processing patterns, data quality, and schema evolution. Workflow orchestration at scale - designing and operating multi-step automated pipelines with retries, observability, and graceful failure handling. Distributed systems fundamentals - you understand how things break at scale: eventual consistency, idempotency, backpressure, job scheduling, and failure modes in distributed compute and storage. Cloud infrastructure fluency - you have shipped and operated real systems on a major cloud provider; you think about cost, reliability, and security as first-class concerns, not afterthoughts. Container orchestration - deploying and operating workloads on Kubernetes at a level where you can debug scheduling issues, design resource allocation, and reason about cluster health without guidance. Full-stack range - comfortable building both the backend and the frontend of an internal product; you can own a feature from database schema to UI without handing off. Production ownership mindset - you've been on-call, triaged incidents under pressure, and improved systems after postmortems. You take reliability personally. Nice to Have ML infrastructure or MLOps experience - understanding of how training jobs run, how model artifacts are managed, and what makes an evaluation pipeline trustworthy; you've worked alongside or directly supported ML researchers. Distributed compute frameworks - experience with large-scale parallel data processing, whether for data transformation, model training, or evaluation. Domain knowledge in robotics or embodied AI - familiarity with robot data formats, sensor telemetry, or the sim-to-real evaluation loop is a significant head start. BI and data warehouse experience - building data models and dashboards that translate raw operational data into decisions for non-technical stakeholders. Dual-cloud or multi-cloud storage - experience reasoning about cost, latency, and consistency tradeoffs across storage providers. Frontend product sense - beyond just shipping features, you have opinions about what makes an internal tool actually usable by non-engineers. What We Offer Competitive equity: stock options with meaningful upside as we scale. 30+ paid days off, including 23 days of annual leave, all UK bank holidays, and additional company closure days (including Christmas-New Year shutdown). Private healthcare, including virtual and in-person care. Pension scheme with 8% total contribution (5% employee, 3% employer) on full earnings. Free daily breakfast, catered lunch, and snacks in-office. Work at the frontier - collaborate daily with world-class engineers, researchers, and product experts building the next generation of AI and humanoid robotics. Real ownership - direct access to founding leadership, meaningful input on product direction, and the ability to drive key initiatives from day one.
18/07/2026
Full time
Here at Humanoid, we believe in a future where robots amplify human potential. That's why we've set out on a mission to build the world's most capable, commercially-scalable, and safe humanoid robots. We're bringing that mission to life with HMND 01 Alpha - our rapidly developed humanoid platform now running in real industrial pilots - and we're growing the team to take it even further. About the Role We're hiring a Sr./Staff Software Engineer to join our Data Engineering team based in London. As the Data Engineering Lead, we strive to create the world's leading, commercially scalable, safe, and advanced humanoid robots that seamlessly integrate into daily life and amplify human capacity. We're looking for experienced Engineers to help build our data platform from the ground up. You'll have the opportunity to define its architecture, make key decisions, and shape how we handle and process petabyte-scale data in the near future. What You'll Do Build the Capability Factory - an internal platform designed for everyone from software engineers to non-technical operators, enabling the entire organization to teach HMND robots new skills at scale, from raw data all the way to deployed capabilities. Curate, preprocess, and manage large-scale datasets for humanoid robot training - a corpus of robot telemetry growing toward petabyte scale. Design and operate highly scalable data pipelines and the compute infrastructure that powers them, ensuring reliability and throughput as data volume and team demands grow. Ensure the quality, accuracy, and consistency of training data across multiple concurrent projects and robot platforms. Collaborate with machine learning teams to shape the Capability Factory, streamline MLOps, and build the evaluation workflows that close the loop between training runs and real-world robot performance. Build data warehouse solutions and BI dashboards that give stakeholders across the organization clear visibility into data collection, model progress, and operational health. Establish and uphold best practices for data management - versioning, access control, security, and compliance. What We're Looking For 5+ years of software engineering experience, with a track record of owning and delivering complex systems end-to-end, not just contributing to them. Strong backend engineering - designing and operating production-grade APIs and services: clean data modeling, reliable error handling, performance under load. Data engineering at PB+ scale - building and maintaining pipelines that move, transform, and validate large volumes of data reliably; understanding of batch and streaming processing patterns, data quality, and schema evolution. Workflow orchestration at scale - designing and operating multi-step automated pipelines with retries, observability, and graceful failure handling. Distributed systems fundamentals - you understand how things break at scale: eventual consistency, idempotency, backpressure, job scheduling, and failure modes in distributed compute and storage. Cloud infrastructure fluency - you have shipped and operated real systems on a major cloud provider; you think about cost, reliability, and security as first-class concerns, not afterthoughts. Container orchestration - deploying and operating workloads on Kubernetes at a level where you can debug scheduling issues, design resource allocation, and reason about cluster health without guidance. Full-stack range - comfortable building both the backend and the frontend of an internal product; you can own a feature from database schema to UI without handing off. Production ownership mindset - you've been on-call, triaged incidents under pressure, and improved systems after postmortems. You take reliability personally. Nice to Have ML infrastructure or MLOps experience - understanding of how training jobs run, how model artifacts are managed, and what makes an evaluation pipeline trustworthy; you've worked alongside or directly supported ML researchers. Distributed compute frameworks - experience with large-scale parallel data processing, whether for data transformation, model training, or evaluation. Domain knowledge in robotics or embodied AI - familiarity with robot data formats, sensor telemetry, or the sim-to-real evaluation loop is a significant head start. BI and data warehouse experience - building data models and dashboards that translate raw operational data into decisions for non-technical stakeholders. Dual-cloud or multi-cloud storage - experience reasoning about cost, latency, and consistency tradeoffs across storage providers. Frontend product sense - beyond just shipping features, you have opinions about what makes an internal tool actually usable by non-engineers. What We Offer Competitive equity: stock options with meaningful upside as we scale. 30+ paid days off, including 23 days of annual leave, all UK bank holidays, and additional company closure days (including Christmas-New Year shutdown). Private healthcare, including virtual and in-person care. Pension scheme with 8% total contribution (5% employee, 3% employer) on full earnings. Free daily breakfast, catered lunch, and snacks in-office. Work at the frontier - collaborate daily with world-class engineers, researchers, and product experts building the next generation of AI and humanoid robotics. Real ownership - direct access to founding leadership, meaningful input on product direction, and the ability to drive key initiatives from day one.
Sr./Staff Software Engineer (Data Platform)
Groupe-Ebra-1
Here at Humanoid, we believe in a future where robots amplify human potential. That's why we've set out on a mission to build the world's most capable, commercially-scalable, and safe humanoid robots. We're bringing that mission to life with HMND 01 Alpha - our rapidly developed humanoid platform now running in real industrial pilots - and we're growing the team to take it even further. About the Role We're hiring a Sr./Staff Software Engineer to join our Data Engineering team based in London. As the Data Engineering Lead, we strive to create the world's leading, commercially scalable, safe, and advanced humanoid robots that seamlessly integrate into daily life and amplify human capacity. We're looking for experienced Engineers to help build our data platform from the ground up. You'll have the opportunity to define its architecture, make key decisions, and shape how we handle and process petabyte-scale data in the near future. What You'll Do Build the Capability Factory - an internal platform designed for everyone from software engineers to non-technical operators, enabling the entire organization to teach HMND robots new skills at scale, from raw data all the way to deployed capabilities. Curate, preprocess, and manage large-scale datasets for humanoid robot training - a corpus of robot telemetry growing toward petabyte scale. Design and operate highly scalable data pipelines and the compute infrastructure that powers them, ensuring reliability and throughput as data volume and team demands grow. Ensure the quality, accuracy, and consistency of training data across multiple concurrent projects and robot platforms. Collaborate with machine learning teams to shape the Capability Factory, streamline MLOps, and build the evaluation workflows that close the loop between training runs and real-world robot performance. Build data warehouse solutions and BI dashboards that give stakeholders across the organization clear visibility into data collection, model progress, and operational health. Establish and uphold best practices for data management - versioning, access control, security, and compliance. What We're Looking For 5+ years of software engineering experience, with a track record of owning and delivering complex systems end-to-end, not just contributing to them. Strong backend engineering - designing and operating production-grade APIs and services: clean data modeling, reliable error handling, performance under load. Data engineering at PB+ scale - building and maintaining pipelines that move, transform, and validate large volumes of data reliably; understanding of batch and streaming processing patterns, data quality, and schema evolution. Workflow orchestration at scale - designing and operating multi-step automated pipelines with retries, observability, and graceful failure handling. Distributed systems fundamentals - you understand how things break at scale: eventual consistency, idempotency, backpressure, job scheduling, and failure modes in distributed compute and storage. Cloud infrastructure fluency - you have shipped and operated real systems on a major cloud provider; you think about cost, reliability, and security as first-class concerns, not afterthoughts. Container orchestration - deploying and operating workloads on Kubernetes at a level where you can debug scheduling issues, design resource allocation, and reason about cluster health without guidance. Full-stack range - comfortable building both the backend and the frontend of an internal product; you can own a feature from database schema to UI without handing off. Production ownership mindset - you've been on-call, triaged incidents under pressure, and improved systems after postmortems. You take reliability personally. Nice to Have ML infrastructure or MLOps experience - understanding of how training jobs run, how model artifacts are managed, and what makes an evaluation pipeline trustworthy; you've worked alongside or directly supported ML researchers. Distributed compute frameworks - experience with large-scale parallel data processing, whether for data transformation, model training, or evaluation. Domain knowledge in robotics or embodied AI - familiarity with robot data formats, sensor telemetry, or the sim-to-real evaluation loop is a significant head start. BI and data warehouse experience - building data models and dashboards that translate raw operational data into decisions for non-technical stakeholders. Dual-cloud or multi-cloud storage - experience reasoning about cost, latency, and consistency tradeoffs across storage providers. Frontend product sense - beyond just shipping features, you have opinions about what makes an internal tool actually usable by non-engineers. What We Offer Competitive equity: stock options with meaningful upside as we scale. 30+ paid days off, including 23 days of annual leave, all UK bank holidays, and additional company closure days (including Christmas-New Year shutdown). Private healthcare, including virtual and in-person care. Pension scheme with 8% total contribution (5% employee, 3% employer) on full earnings. Free daily breakfast, catered lunch, and snacks in-office. Work at the frontier - collaborate daily with world-class engineers, researchers, and product experts building the next generation of AI and humanoid robotics. Real ownership - direct access to founding leadership, meaningful input on product direction, and the ability to drive key initiatives from day one.
17/07/2026
Full time
Here at Humanoid, we believe in a future where robots amplify human potential. That's why we've set out on a mission to build the world's most capable, commercially-scalable, and safe humanoid robots. We're bringing that mission to life with HMND 01 Alpha - our rapidly developed humanoid platform now running in real industrial pilots - and we're growing the team to take it even further. About the Role We're hiring a Sr./Staff Software Engineer to join our Data Engineering team based in London. As the Data Engineering Lead, we strive to create the world's leading, commercially scalable, safe, and advanced humanoid robots that seamlessly integrate into daily life and amplify human capacity. We're looking for experienced Engineers to help build our data platform from the ground up. You'll have the opportunity to define its architecture, make key decisions, and shape how we handle and process petabyte-scale data in the near future. What You'll Do Build the Capability Factory - an internal platform designed for everyone from software engineers to non-technical operators, enabling the entire organization to teach HMND robots new skills at scale, from raw data all the way to deployed capabilities. Curate, preprocess, and manage large-scale datasets for humanoid robot training - a corpus of robot telemetry growing toward petabyte scale. Design and operate highly scalable data pipelines and the compute infrastructure that powers them, ensuring reliability and throughput as data volume and team demands grow. Ensure the quality, accuracy, and consistency of training data across multiple concurrent projects and robot platforms. Collaborate with machine learning teams to shape the Capability Factory, streamline MLOps, and build the evaluation workflows that close the loop between training runs and real-world robot performance. Build data warehouse solutions and BI dashboards that give stakeholders across the organization clear visibility into data collection, model progress, and operational health. Establish and uphold best practices for data management - versioning, access control, security, and compliance. What We're Looking For 5+ years of software engineering experience, with a track record of owning and delivering complex systems end-to-end, not just contributing to them. Strong backend engineering - designing and operating production-grade APIs and services: clean data modeling, reliable error handling, performance under load. Data engineering at PB+ scale - building and maintaining pipelines that move, transform, and validate large volumes of data reliably; understanding of batch and streaming processing patterns, data quality, and schema evolution. Workflow orchestration at scale - designing and operating multi-step automated pipelines with retries, observability, and graceful failure handling. Distributed systems fundamentals - you understand how things break at scale: eventual consistency, idempotency, backpressure, job scheduling, and failure modes in distributed compute and storage. Cloud infrastructure fluency - you have shipped and operated real systems on a major cloud provider; you think about cost, reliability, and security as first-class concerns, not afterthoughts. Container orchestration - deploying and operating workloads on Kubernetes at a level where you can debug scheduling issues, design resource allocation, and reason about cluster health without guidance. Full-stack range - comfortable building both the backend and the frontend of an internal product; you can own a feature from database schema to UI without handing off. Production ownership mindset - you've been on-call, triaged incidents under pressure, and improved systems after postmortems. You take reliability personally. Nice to Have ML infrastructure or MLOps experience - understanding of how training jobs run, how model artifacts are managed, and what makes an evaluation pipeline trustworthy; you've worked alongside or directly supported ML researchers. Distributed compute frameworks - experience with large-scale parallel data processing, whether for data transformation, model training, or evaluation. Domain knowledge in robotics or embodied AI - familiarity with robot data formats, sensor telemetry, or the sim-to-real evaluation loop is a significant head start. BI and data warehouse experience - building data models and dashboards that translate raw operational data into decisions for non-technical stakeholders. Dual-cloud or multi-cloud storage - experience reasoning about cost, latency, and consistency tradeoffs across storage providers. Frontend product sense - beyond just shipping features, you have opinions about what makes an internal tool actually usable by non-engineers. What We Offer Competitive equity: stock options with meaningful upside as we scale. 30+ paid days off, including 23 days of annual leave, all UK bank holidays, and additional company closure days (including Christmas-New Year shutdown). Private healthcare, including virtual and in-person care. Pension scheme with 8% total contribution (5% employee, 3% employer) on full earnings. Free daily breakfast, catered lunch, and snacks in-office. Work at the frontier - collaborate daily with world-class engineers, researchers, and product experts building the next generation of AI and humanoid robotics. Real ownership - direct access to founding leadership, meaningful input on product direction, and the ability to drive key initiatives from day one.
Software Engineer, General
United States Digital Space LLC
About the companythe company is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, the company takes them from concept to production on a single, unified platform. We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters. We're a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we're building the next generation of enterprise AI and we're looking for exceptional people to help us scale. Who You Are You're a solid engineer with 2-4 years of experience building backend systems and platform infrastructure. You write clean, well-abstracted code with proper design patterns and comprehensive test coverage. You're comfortable working on both the Compute Platform (multi-cloud orchestration, resource management) and Inference Platform (model serving, autoscaling) under the guidance of senior engineers and platform leads. You have strong proficiency in Golang and understand how to build maintainable, production-grade distributed systems. You take pride in code quality, enjoy collaborating on low-level designs, and are eager to learn from experienced engineers while contributing meaningfully to critical infrastructure components. You're product minded, you understand how your technical decisions impact developers using the company's platform and think about the end to end user experience. You're a team player comfortable wearing multiple hats one day you're building product features, the next you're joining customer calls to understand their deployment challenges, and the day after you're helping with UI/UX, customer success, documentation and product ops. What You'll Do Platform Development & Implementation Build and maintain platform services across the company's Compute and Inference platforms, working closely with senior engineers and platform leads Implement features for multi cloud orchestration, resource scheduling, model deployment pipelines, and autoscaling systems Write well maintained, production grade code with proper abstractions, design patterns, and comprehensive test coverage Contribute to low level design (LLD) including service APIs, database schema design, data models, and component interactions Collaborate with senior engineers on high level design discussions, providing implementation perspectives and feasibility inputs Backend Systems & Distributed Infrastructure Develop RESTful APIs and gRPC services for platform control planes, resource management, and inference serving Design and implement database schemas for storing platform state, resource metadata, billing data, and observability metrics Work with distributed storage systems, message queues (Kafka, RabbitMQ), and databases (PostgreSQL, Redis) to build reliable platform components Build event driven architectures for asynchronous processing, job scheduling, and platform automation Implement monitoring, logging, and alerting for platform services to ensure production reliability Code Quality & Engineering Excellence Write comprehensive unit tests, integration tests, and end to end tests to ensure code reliability Participate in code reviews, providing constructive feedback and learning from senior engineers' perspectives Refactor existing code to improve maintainability, performance, and scalability Document design decisions, API specifications, and operational runbooks for platform services Debug production issues and contribute to incident response and post mortems. Requirements Technical Skills & Experience 2-4 years of experience in backend engineering, platform development, or distributed systems Strong proficiency in Golang you write idiomatic Go code with proper error handling, concurrency patterns, and testing Solid understanding of backend systems fundamentals: RESTful APIs, microservices architecture, and API design principles Hands on experience with databases (PostgreSQL, MySQL) including schema design, query optimization, and transactions Familiarity with storage systems (object storage like S3, block storage, distributed file systems) and their use cases Experience working with message queues (Kafka, RabbitMQ, NATS) and event driven architectures Understanding of distributed systems concepts: consensus, eventual consistency, fault tolerance, and retry mechanisms Experience with containerization (Docker) and basic Kubernetes concepts Knowledge of testing frameworks and practices (unit tests, integration tests, mocking) Familiarity with Git, CI/CD pipelines, and modern development workflows Exposure to cloud platforms (AWS/GCP/Azure) and their core services is a plus Experience with infrastructure as code (Terraform) or observability tools (Prometheus, Grafana) is beneficial Bonus/ Good to Have HPC & Cluster Management: Experience handling large-scale HPC clusters using Kubernetes and Slurm for job scheduling, resource allocation, and workload orchestration Data Engineering: Expertise with data pipelines, ETL systems, and large-scale data processing frameworks Systems Level Programming: Experience with low-level systems programming such as storage systems, Kubernetes operators, OS-level software development, or daemon services (llm d, system agents) ML Platform Engineering: Experience productionizing ML pipelines, batch job orchestration, model fine tuning workflows, and Jupyter notebook orchestration systems Enterprise Deployment: Experience platformizing and packaging software for on premises deployments or customer VPC installations with emphasis on security, compliance, and operational simplicity Benefits Preferred Attributes High ownership, self driven and biased for action. Strong strategic thinking and ability to connect technical decisions to business impact. Excellent communication and mentoring skills. Thrives in ambiguity, fast paced environments, and early stage startup culture. Why Join the company? Work directly with high pedigree founders shaping technical and product strategy. Build infrastructure powering the future of AI computers globally. Significant ownership and impact with equity reflective of your contributions. Competitive compensation, flexible work options, and wellness benefits.
16/07/2026
Full time
About the companythe company is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, the company takes them from concept to production on a single, unified platform. We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters. We're a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we're building the next generation of enterprise AI and we're looking for exceptional people to help us scale. Who You Are You're a solid engineer with 2-4 years of experience building backend systems and platform infrastructure. You write clean, well-abstracted code with proper design patterns and comprehensive test coverage. You're comfortable working on both the Compute Platform (multi-cloud orchestration, resource management) and Inference Platform (model serving, autoscaling) under the guidance of senior engineers and platform leads. You have strong proficiency in Golang and understand how to build maintainable, production-grade distributed systems. You take pride in code quality, enjoy collaborating on low-level designs, and are eager to learn from experienced engineers while contributing meaningfully to critical infrastructure components. You're product minded, you understand how your technical decisions impact developers using the company's platform and think about the end to end user experience. You're a team player comfortable wearing multiple hats one day you're building product features, the next you're joining customer calls to understand their deployment challenges, and the day after you're helping with UI/UX, customer success, documentation and product ops. What You'll Do Platform Development & Implementation Build and maintain platform services across the company's Compute and Inference platforms, working closely with senior engineers and platform leads Implement features for multi cloud orchestration, resource scheduling, model deployment pipelines, and autoscaling systems Write well maintained, production grade code with proper abstractions, design patterns, and comprehensive test coverage Contribute to low level design (LLD) including service APIs, database schema design, data models, and component interactions Collaborate with senior engineers on high level design discussions, providing implementation perspectives and feasibility inputs Backend Systems & Distributed Infrastructure Develop RESTful APIs and gRPC services for platform control planes, resource management, and inference serving Design and implement database schemas for storing platform state, resource metadata, billing data, and observability metrics Work with distributed storage systems, message queues (Kafka, RabbitMQ), and databases (PostgreSQL, Redis) to build reliable platform components Build event driven architectures for asynchronous processing, job scheduling, and platform automation Implement monitoring, logging, and alerting for platform services to ensure production reliability Code Quality & Engineering Excellence Write comprehensive unit tests, integration tests, and end to end tests to ensure code reliability Participate in code reviews, providing constructive feedback and learning from senior engineers' perspectives Refactor existing code to improve maintainability, performance, and scalability Document design decisions, API specifications, and operational runbooks for platform services Debug production issues and contribute to incident response and post mortems. Requirements Technical Skills & Experience 2-4 years of experience in backend engineering, platform development, or distributed systems Strong proficiency in Golang you write idiomatic Go code with proper error handling, concurrency patterns, and testing Solid understanding of backend systems fundamentals: RESTful APIs, microservices architecture, and API design principles Hands on experience with databases (PostgreSQL, MySQL) including schema design, query optimization, and transactions Familiarity with storage systems (object storage like S3, block storage, distributed file systems) and their use cases Experience working with message queues (Kafka, RabbitMQ, NATS) and event driven architectures Understanding of distributed systems concepts: consensus, eventual consistency, fault tolerance, and retry mechanisms Experience with containerization (Docker) and basic Kubernetes concepts Knowledge of testing frameworks and practices (unit tests, integration tests, mocking) Familiarity with Git, CI/CD pipelines, and modern development workflows Exposure to cloud platforms (AWS/GCP/Azure) and their core services is a plus Experience with infrastructure as code (Terraform) or observability tools (Prometheus, Grafana) is beneficial Bonus/ Good to Have HPC & Cluster Management: Experience handling large-scale HPC clusters using Kubernetes and Slurm for job scheduling, resource allocation, and workload orchestration Data Engineering: Expertise with data pipelines, ETL systems, and large-scale data processing frameworks Systems Level Programming: Experience with low-level systems programming such as storage systems, Kubernetes operators, OS-level software development, or daemon services (llm d, system agents) ML Platform Engineering: Experience productionizing ML pipelines, batch job orchestration, model fine tuning workflows, and Jupyter notebook orchestration systems Enterprise Deployment: Experience platformizing and packaging software for on premises deployments or customer VPC installations with emphasis on security, compliance, and operational simplicity Benefits Preferred Attributes High ownership, self driven and biased for action. Strong strategic thinking and ability to connect technical decisions to business impact. Excellent communication and mentoring skills. Thrives in ambiguity, fast paced environments, and early stage startup culture. Why Join the company? Work directly with high pedigree founders shaping technical and product strategy. Build infrastructure powering the future of AI computers globally. Significant ownership and impact with equity reflective of your contributions. Competitive compensation, flexible work options, and wellness benefits.
Senior Software Engineer, Inference Platform
United States Digital Space LLC
About the companythe company is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, the company takes them from concept to production on a single, unified platform. We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters. We're a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we're building the next generation of enterprise AI and we're looking for exceptional people to help us scale. Who You Are You're a seasoned engineer who has built and scaled high-performance inference systems for AI/ML workloads. You understand the complexities of serving models at scale latency optimization, resource orchestration, autoscaling dynamics, and production reliability. You've designed distributed systems that handle thousands of requests per second while maintaining sub second response times and cost efficiency. Experience with Golang is strongly preferred, and exposure to inference engines (vLLM, TGI, TensorRT), containerization, and distributed systems is an added bonus. You take ownership of platform level decisions, think strategically about performance vs. cost trade offs, and want your work to power AI inference for thousands of developers globally. You're product minded, you understand how your technical decisions impact developers using the company's platform and think about the end to end user experience. You're a team player comfortable wearing multiple hats one day you're optimizing inference latency, the next you're joining customer calls to understand their deployment challenges, and the day after you're helping with UI/UX, customer success, documentation and product ops. What You'll Do Inference Platform Architecture & Core Services Design and build the company's inference service platform the backbone for serving AI models at scale across diverse workloads Own and architect core platform components: AI Gateway, Resource Orchestrator, Runtime Engines, and Autoscaler Design highly modular, scalable, and extensible low level designs (LLDs) for inference infrastructure components Lead high level design discussions, establish architectural patterns, and drive technical decision making for the inference stack Model Deployment & Lifecycle Management Understand and optimize the dynamics of model deployment, version upgrades, and rollback strategies Build robust deployment pipelines for seamless model updates with zero downtime deployments Design intelligent routing systems for multi model serving, A/B testing, and canary deployments Implement strategies for efficient GPU utilization and model cold start optimization Performance & Distributed Systems Implement highly performant and optimized software for low latency, high throughput inference serving Build and debug production grade code in distributed systems handling real time AI workloads Optimize inference pipelines for latency, throughput, batching efficiency, and resource utilization Design fault tolerant systems with graceful degradation and automatic recovery mechanisms Observability & Engineering Excellence Build high performance telemetry and observability stack for inference metrics, performance tracking, and debugging Implement comprehensive monitoring for model latency, throughput, error rates, GPU utilization, and cost per inference Conduct thorough code reviews to maintain code quality, performance standards, and architectural consistency Establish engineering best practices for testing, documentation, and production readiness. Requirements Technical Skills & Experience 4+ years of experience building and scaling backend systems, distributed platforms, or inference infrastructure Strong understanding of AI/ML inference systems and experience with inference engines (vLLM, TGI, TensorRT LLM, or similar) Deep knowledge of distributed systems design, microservices architecture, and API gateway patterns Proficiency in Golang strongly preferred; Python, Rust, C++ for performance critical components a plus Experience with container orchestration (Kubernetes, Docker) and infrastructure as code Solid understanding of autoscaling strategies, load balancing, and resource scheduling algorithms Experience building high throughput, low latency systems with sub 100ms response time requirements Familiarity with message queues (Kafka, RabbitMQ), databases (PostgreSQL, Redis), and event driven architectures Knowledge of GPU computing, model serving optimizations (batching, quantization, multi tenancy), and resource allocation Experience with observability tools (Prometheus, Grafana, OpenTelemetry) and distributed tracing Understanding of API design, rate limiting, authentication/authorization, and security best practices Exposure to AI model deployment workflows and model lifecycle management is highly desirable Bonus / Good to Have HPC & Cluster Management: Experience handling large scale HPC clusters using Kubernetes and Slurm for job scheduling, resource allocation, and workload orchestration Data Engineering: Expertise with data pipelines, ETL systems, and large scale data processing frameworks Systems-Level Programming: Experience with low level systems programming such as storage systems, Kubernetes operators, OS-level software development, or daemon services (llm d, system agents) ML Platform Engineering: Experience productionizing ML pipelines, batch job orchestration, model fine tuning workflows, and Jupyter notebook orchestration systems Enterprise Deployment: Experience platformizing and packaging software for on premises deployments or customer VPC installations with emphasis on security, compliance, and operational simplicity Benefits Preferred Attributes High ownership, self driven and biased for action. Strong strategic thinking and ability to connect technical decisions to business impact. Excellent communication and mentoring skills. Thrives in ambiguity, fast paced environments, and early-stage startup culture. Why Join the company? Work directly with high pedigree founders shaping technical and product strategy. Build infrastructure powering the future of AI computers globally. Significant ownership and impact with equity reflective of your contributions. Competitive compensation, flexible work options, and wellness benefits
16/07/2026
Full time
About the companythe company is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, the company takes them from concept to production on a single, unified platform. We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters. We're a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we're building the next generation of enterprise AI and we're looking for exceptional people to help us scale. Who You Are You're a seasoned engineer who has built and scaled high-performance inference systems for AI/ML workloads. You understand the complexities of serving models at scale latency optimization, resource orchestration, autoscaling dynamics, and production reliability. You've designed distributed systems that handle thousands of requests per second while maintaining sub second response times and cost efficiency. Experience with Golang is strongly preferred, and exposure to inference engines (vLLM, TGI, TensorRT), containerization, and distributed systems is an added bonus. You take ownership of platform level decisions, think strategically about performance vs. cost trade offs, and want your work to power AI inference for thousands of developers globally. You're product minded, you understand how your technical decisions impact developers using the company's platform and think about the end to end user experience. You're a team player comfortable wearing multiple hats one day you're optimizing inference latency, the next you're joining customer calls to understand their deployment challenges, and the day after you're helping with UI/UX, customer success, documentation and product ops. What You'll Do Inference Platform Architecture & Core Services Design and build the company's inference service platform the backbone for serving AI models at scale across diverse workloads Own and architect core platform components: AI Gateway, Resource Orchestrator, Runtime Engines, and Autoscaler Design highly modular, scalable, and extensible low level designs (LLDs) for inference infrastructure components Lead high level design discussions, establish architectural patterns, and drive technical decision making for the inference stack Model Deployment & Lifecycle Management Understand and optimize the dynamics of model deployment, version upgrades, and rollback strategies Build robust deployment pipelines for seamless model updates with zero downtime deployments Design intelligent routing systems for multi model serving, A/B testing, and canary deployments Implement strategies for efficient GPU utilization and model cold start optimization Performance & Distributed Systems Implement highly performant and optimized software for low latency, high throughput inference serving Build and debug production grade code in distributed systems handling real time AI workloads Optimize inference pipelines for latency, throughput, batching efficiency, and resource utilization Design fault tolerant systems with graceful degradation and automatic recovery mechanisms Observability & Engineering Excellence Build high performance telemetry and observability stack for inference metrics, performance tracking, and debugging Implement comprehensive monitoring for model latency, throughput, error rates, GPU utilization, and cost per inference Conduct thorough code reviews to maintain code quality, performance standards, and architectural consistency Establish engineering best practices for testing, documentation, and production readiness. Requirements Technical Skills & Experience 4+ years of experience building and scaling backend systems, distributed platforms, or inference infrastructure Strong understanding of AI/ML inference systems and experience with inference engines (vLLM, TGI, TensorRT LLM, or similar) Deep knowledge of distributed systems design, microservices architecture, and API gateway patterns Proficiency in Golang strongly preferred; Python, Rust, C++ for performance critical components a plus Experience with container orchestration (Kubernetes, Docker) and infrastructure as code Solid understanding of autoscaling strategies, load balancing, and resource scheduling algorithms Experience building high throughput, low latency systems with sub 100ms response time requirements Familiarity with message queues (Kafka, RabbitMQ), databases (PostgreSQL, Redis), and event driven architectures Knowledge of GPU computing, model serving optimizations (batching, quantization, multi tenancy), and resource allocation Experience with observability tools (Prometheus, Grafana, OpenTelemetry) and distributed tracing Understanding of API design, rate limiting, authentication/authorization, and security best practices Exposure to AI model deployment workflows and model lifecycle management is highly desirable Bonus / Good to Have HPC & Cluster Management: Experience handling large scale HPC clusters using Kubernetes and Slurm for job scheduling, resource allocation, and workload orchestration Data Engineering: Expertise with data pipelines, ETL systems, and large scale data processing frameworks Systems-Level Programming: Experience with low level systems programming such as storage systems, Kubernetes operators, OS-level software development, or daemon services (llm d, system agents) ML Platform Engineering: Experience productionizing ML pipelines, batch job orchestration, model fine tuning workflows, and Jupyter notebook orchestration systems Enterprise Deployment: Experience platformizing and packaging software for on premises deployments or customer VPC installations with emphasis on security, compliance, and operational simplicity Benefits Preferred Attributes High ownership, self driven and biased for action. Strong strategic thinking and ability to connect technical decisions to business impact. Excellent communication and mentoring skills. Thrives in ambiguity, fast paced environments, and early-stage startup culture. Why Join the company? Work directly with high pedigree founders shaping technical and product strategy. Build infrastructure powering the future of AI computers globally. Significant ownership and impact with equity reflective of your contributions. Competitive compensation, flexible work options, and wellness benefits
Member of Technical Staff (AI Infrastructure Engineer)
Pantera Capital
Location London Employment Type Full time Location Type Hybrid Department AI We are looking for an AI Infra engineer to join our growing team. We work with Kubernetes, Slurm, Python, C++, PyTorch, and primarily on AWS. As an AI Infrastructure Engineer, you will be partnering closely with our Inference and Research teams to build, deploy, and optimize our large-scale AI training and inference clusters. Responsibilities Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads Manage and optimize Slurm-based HPC environments for distributed training of large language models Develop robust APIs and orchestration systems for both training pipelines and inference services Implement resource scheduling and job management systems across heterogeneous compute environments Benchmark system performance, diagnose bottlenecks, and implement improvements across both training and inference infrastructure Build monitoring, alerting, and observability solutions tailored to ML workloads running on Kubernetes and Slurm Respond swiftly to system outages and collaborate across teams to maintain high uptime for critical training runs and inference services Optimize cluster utilization and implement autoscaling strategies for dynamic workload demands Qualifications Strong expertise in Kubernetes administration, including custom resource definitions, operators, and cluster management Hands-on experience with Slurm workload management, including job scheduling, resource allocation, and cluster optimization Experience with deploying and managing distributed training systems at scale Deep understanding of container orchestration and distributed systems architecture High level familiarity with LLM architecture and training processes (Multi-Head Attention, Multi/Grouped-Query, distributed training strategies) Experience managing GPU clusters and optimizing compute resource utilization Required Skills Expert-level Kubernetes administration and YAML configuration management Proficiency with Slurm job scheduling, resource management, and cluster configuration Python and C++ programming with focus on systems and infrastructure automation Hands-on experience with ML frameworks such as PyTorch in distributed training contexts Strong understanding of networking, storage, and compute resource management for ML workloads Experience developing APIs and managing distributed systems for both batch and real-time workloads Solid debugging and monitoring skills with expertise in observability tools for containerized environments Preferred Skills Experience with Kubernetes operators and custom controllers for ML workloads Advanced Slurm administration including multi-cluster federation and advanced scheduling policies Familiarity with GPU cluster management and CUDA optimization Experience with other ML frameworks like TensorFlow or distributed training libraries Background in HPC environments, parallel computing, and high-performance networking Knowledge of infrastructure as code (Terraform, Ansible) and GitOps practices Experience with container registries, image optimization, and multi-stage builds for ML workloads Required Experience Demonstrated experience managing large-scale Kubernetes deployments in production environments Proven track record with Slurm cluster administration and HPC workload management Previous roles in SRE, DevOps, or Platform Engineering with focus on ML infrastructure Experience supporting both long-running training jobs and high-availability inference services Ideally, 3-5 years of relevant experience in ML systems deployment with specific focus on cluster orchestration and resource management
14/07/2026
Full time
Location London Employment Type Full time Location Type Hybrid Department AI We are looking for an AI Infra engineer to join our growing team. We work with Kubernetes, Slurm, Python, C++, PyTorch, and primarily on AWS. As an AI Infrastructure Engineer, you will be partnering closely with our Inference and Research teams to build, deploy, and optimize our large-scale AI training and inference clusters. Responsibilities Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads Manage and optimize Slurm-based HPC environments for distributed training of large language models Develop robust APIs and orchestration systems for both training pipelines and inference services Implement resource scheduling and job management systems across heterogeneous compute environments Benchmark system performance, diagnose bottlenecks, and implement improvements across both training and inference infrastructure Build monitoring, alerting, and observability solutions tailored to ML workloads running on Kubernetes and Slurm Respond swiftly to system outages and collaborate across teams to maintain high uptime for critical training runs and inference services Optimize cluster utilization and implement autoscaling strategies for dynamic workload demands Qualifications Strong expertise in Kubernetes administration, including custom resource definitions, operators, and cluster management Hands-on experience with Slurm workload management, including job scheduling, resource allocation, and cluster optimization Experience with deploying and managing distributed training systems at scale Deep understanding of container orchestration and distributed systems architecture High level familiarity with LLM architecture and training processes (Multi-Head Attention, Multi/Grouped-Query, distributed training strategies) Experience managing GPU clusters and optimizing compute resource utilization Required Skills Expert-level Kubernetes administration and YAML configuration management Proficiency with Slurm job scheduling, resource management, and cluster configuration Python and C++ programming with focus on systems and infrastructure automation Hands-on experience with ML frameworks such as PyTorch in distributed training contexts Strong understanding of networking, storage, and compute resource management for ML workloads Experience developing APIs and managing distributed systems for both batch and real-time workloads Solid debugging and monitoring skills with expertise in observability tools for containerized environments Preferred Skills Experience with Kubernetes operators and custom controllers for ML workloads Advanced Slurm administration including multi-cluster federation and advanced scheduling policies Familiarity with GPU cluster management and CUDA optimization Experience with other ML frameworks like TensorFlow or distributed training libraries Background in HPC environments, parallel computing, and high-performance networking Knowledge of infrastructure as code (Terraform, Ansible) and GitOps practices Experience with container registries, image optimization, and multi-stage builds for ML workloads Required Experience Demonstrated experience managing large-scale Kubernetes deployments in production environments Proven track record with Slurm cluster administration and HPC workload management Previous roles in SRE, DevOps, or Platform Engineering with focus on ML infrastructure Experience supporting both long-running training jobs and high-availability inference services Ideally, 3-5 years of relevant experience in ML systems deployment with specific focus on cluster orchestration and resource management
Sr./Staff Software Engineer (Data Team)
Humanoid
Here at Humanoid, we believe in a future where robots amplify human potential. That's why we've set out on a mission to build the world's most capable, commercially-scalable, and safe humanoid robots. We're bringing that mission to life with HMND 01 Alpha - our rapidly developed humanoid platform now running in real industrial pilots - and we're growing the team to take it even further. About the Role We're hiring a Sr./Staff Software Engineer to join our Data Engineering team based in London. As the Data Engineering Lead, we strive to create the world's leading, commercially scalable, safe, and advanced humanoid robots that seamlessly integrate into daily life and amplify human capacity. We're looking for experienced Engineers to help build our data platform from the ground up. You'll have the opportunity to define its architecture, make key decisions, and shape how we handle and process petabyte-scale data in the near future. What You'll Do Build the Capability Factory - an internal platform designed for everyone from software engineers to non-technical operators, enabling the entire organization to teach HMND robots new skills at scale, from raw data all the way to deployed capabilities. Curate, preprocess, and manage large-scale datasets for humanoid robot training - a corpus of robot telemetry growing toward petabyte scale. Design and operate highly scalable data pipelines and the compute infrastructure that powers them, ensuring reliability and throughput as data volume and team demands grow. Ensure the quality, accuracy, and consistency of training data across multiple concurrent projects and robot platforms. Collaborate with machine learning teams to shape the Capability Factory, streamline MLOps, and build the evaluation workflows that close the loop between training runs and real-world robot performance. Build data warehouse solutions and BI dashboards that give stakeholders across the organization clear visibility into data collection, model progress, and operational health. Establish and uphold best practices for data management - versioning, access control, security, and compliance. What We're Looking For 5+ years of software engineering experience, with a track record of owning and delivering complex systems end-to-end, not just contributing to them. Strong backend engineering - designing and operating production-grade APIs and services: clean data modeling, reliable error handling, performance under load. Data engineering at PB+ scale - building and maintaining pipelines that move, transform, and validate large volumes of data reliably; understanding of batch and streaming processing patterns, data quality, and schema evolution. Workflow orchestration at scale - designing and operating multi-step automated pipelines with retries, observability, and graceful failure handling. Distributed systems fundamentals - you understand how things break at scale: eventual consistency, idempotency, backpressure, job scheduling, and failure modes in distributed compute and storage. Cloud infrastructure fluency - you have shipped and operated real systems on a major cloud provider; you think about cost, reliability, and security as first-class concerns, not afterthoughts. Container orchestration - deploying and operating workloads on Kubernetes at a level where you can debug scheduling issues, design resource allocation, and reason about cluster health without guidance. Full-stack range - comfortable building both the backend and the frontend of an internal product; you can own a feature from database schema to UI without handing off. Production ownership mindset - you've been on-call, triaged incidents under pressure, and improved systems after postmortems. You take reliability personally. Nice to Have ML infrastructure or MLOps experience - understanding of how training jobs run, how model artifacts are managed, and what makes an evaluation pipeline trustworthy; you've worked alongside or directly supported ML researchers. Distributed compute frameworks - experience with large-scale parallel data processing, whether for data transformation, model training, or evaluation. Domain knowledge in robotics or embodied AI - familiarity with robot data formats, sensor telemetry, or the sim-to-real evaluation loop is a significant head start. BI and data warehouse experience - building data models and dashboards that translate raw operational data into decisions for non-technical stakeholders. Dual-cloud or multi-cloud storage - experience reasoning about cost, latency, and consistency tradeoffs across storage providers. Frontend product sense - beyond just shipping features, you have opinions about what makes an internal tool actually usable by non-engineers. What We Offer Competitive equity: stock options with meaningful upside as we scale. 30+ days off, including 23 days annual leave, all UK bank holidays, and additional company closure days (including Christmas-New Year shutdown). Private healthcare, including virtual and in-person care. Pension scheme with 8% total contribution (5% employee, 3% employer) on full earnings. Free daily breakfast, catered lunch, and snacks in-office. Work at the frontier - collaborate daily with world class engineers, researchers, and product experts building the next generation of AI and humanoid robotics. Real ownership - direct access to founding leadership, meaningful input on product direction, and the ability to drive key initiatives from day one.
12/07/2026
Full time
Here at Humanoid, we believe in a future where robots amplify human potential. That's why we've set out on a mission to build the world's most capable, commercially-scalable, and safe humanoid robots. We're bringing that mission to life with HMND 01 Alpha - our rapidly developed humanoid platform now running in real industrial pilots - and we're growing the team to take it even further. About the Role We're hiring a Sr./Staff Software Engineer to join our Data Engineering team based in London. As the Data Engineering Lead, we strive to create the world's leading, commercially scalable, safe, and advanced humanoid robots that seamlessly integrate into daily life and amplify human capacity. We're looking for experienced Engineers to help build our data platform from the ground up. You'll have the opportunity to define its architecture, make key decisions, and shape how we handle and process petabyte-scale data in the near future. What You'll Do Build the Capability Factory - an internal platform designed for everyone from software engineers to non-technical operators, enabling the entire organization to teach HMND robots new skills at scale, from raw data all the way to deployed capabilities. Curate, preprocess, and manage large-scale datasets for humanoid robot training - a corpus of robot telemetry growing toward petabyte scale. Design and operate highly scalable data pipelines and the compute infrastructure that powers them, ensuring reliability and throughput as data volume and team demands grow. Ensure the quality, accuracy, and consistency of training data across multiple concurrent projects and robot platforms. Collaborate with machine learning teams to shape the Capability Factory, streamline MLOps, and build the evaluation workflows that close the loop between training runs and real-world robot performance. Build data warehouse solutions and BI dashboards that give stakeholders across the organization clear visibility into data collection, model progress, and operational health. Establish and uphold best practices for data management - versioning, access control, security, and compliance. What We're Looking For 5+ years of software engineering experience, with a track record of owning and delivering complex systems end-to-end, not just contributing to them. Strong backend engineering - designing and operating production-grade APIs and services: clean data modeling, reliable error handling, performance under load. Data engineering at PB+ scale - building and maintaining pipelines that move, transform, and validate large volumes of data reliably; understanding of batch and streaming processing patterns, data quality, and schema evolution. Workflow orchestration at scale - designing and operating multi-step automated pipelines with retries, observability, and graceful failure handling. Distributed systems fundamentals - you understand how things break at scale: eventual consistency, idempotency, backpressure, job scheduling, and failure modes in distributed compute and storage. Cloud infrastructure fluency - you have shipped and operated real systems on a major cloud provider; you think about cost, reliability, and security as first-class concerns, not afterthoughts. Container orchestration - deploying and operating workloads on Kubernetes at a level where you can debug scheduling issues, design resource allocation, and reason about cluster health without guidance. Full-stack range - comfortable building both the backend and the frontend of an internal product; you can own a feature from database schema to UI without handing off. Production ownership mindset - you've been on-call, triaged incidents under pressure, and improved systems after postmortems. You take reliability personally. Nice to Have ML infrastructure or MLOps experience - understanding of how training jobs run, how model artifacts are managed, and what makes an evaluation pipeline trustworthy; you've worked alongside or directly supported ML researchers. Distributed compute frameworks - experience with large-scale parallel data processing, whether for data transformation, model training, or evaluation. Domain knowledge in robotics or embodied AI - familiarity with robot data formats, sensor telemetry, or the sim-to-real evaluation loop is a significant head start. BI and data warehouse experience - building data models and dashboards that translate raw operational data into decisions for non-technical stakeholders. Dual-cloud or multi-cloud storage - experience reasoning about cost, latency, and consistency tradeoffs across storage providers. Frontend product sense - beyond just shipping features, you have opinions about what makes an internal tool actually usable by non-engineers. What We Offer Competitive equity: stock options with meaningful upside as we scale. 30+ days off, including 23 days annual leave, all UK bank holidays, and additional company closure days (including Christmas-New Year shutdown). Private healthcare, including virtual and in-person care. Pension scheme with 8% total contribution (5% employee, 3% employer) on full earnings. Free daily breakfast, catered lunch, and snacks in-office. Work at the frontier - collaborate daily with world class engineers, researchers, and product experts building the next generation of AI and humanoid robotics. Real ownership - direct access to founding leadership, meaningful input on product direction, and the ability to drive key initiatives from day one.
Software Engineer, GPU Infrastructure- ChatGPT Engineering
OpenAI
About the Team ChatGPT Engineering builds and operates the compute platform powering one of the world's largest AI products. Every ChatGPT conversation relies on massive GPU clusters serving inference workloads with high reliability, efficiency, and performance. As our GPU fleet continues to grow, we're investing in the infrastructure that operates it. Our team builds the tooling, automation, and intelligent systems that make GPU infrastructure scalable, observable, and increasingly autonomous. We work across production engineering, distributed systems, capacity management, and AI-powered operational tooling to help researchers and product teams move faster while maximizing the efficiency of every GPU. This is a unique opportunity to work on infrastructure at the frontier of AI, where small improvements in fleet efficiency, reliability, and automation have an outsized impact on the development and deployment of AGI. About the Role We're looking for a Software Engineer with deep experience operating large-scale GPU or compute infrastructure. You'll design and build the systems that manage GPU clusters at scale-from fleet health and capacity planning to operational automation and intelligent agents that reduce manual intervention. You'll partner closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and overall compute utilization. This role is ideal for engineers who enjoy solving complex operational challenges, building internal platforms, and working on infrastructure that directly powers frontier AI. In This Role, You Will Design, build, and operate software that manages large-scale GPU infrastructure supporting ChatGPT inference. Build internal platforms, tooling, and AI-powered agents that automate fleet operations and reduce operational overhead. Improve observability, reliability, and operational efficiency across thousands of GPUs. Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response. Identify infrastructure bottlenecks and implement solutions that improve utilization, scalability, and performance. Partner closely with research, platform, networking, and systems teams to continuously improve our compute platform. Help establish engineering best practices around operational excellence, automation, and infrastructure reliability. You Might Thrive in This Role If You Have experience operating large-scale production infrastructure, preferably GPU clusters or other compute-intensive distributed systems. Have a background in Production Engineering, Site Reliability Engineering (SRE), Infrastructure Engineering, or Platform Engineering. Have built software that automates operational workflows rather than relying on manual processes. Have experience with Kubernetes, Linux systems, container orchestration, or distributed infrastructure. Understand infrastructure observability, monitoring, capacity planning, and incident management. Enjoy identifying cross-team pain points and building reusable platforms that improve developer productivity. Are comfortable working across software engineering and systems operations, owning problems end-to-end. Thrive in fast-moving environments with significant technical ambiguity. Qualifications 5+ years of software engineering experience building production infrastructure. Strong programming skills in Go, Python, C++, Rust, or similar systems languages. Experience designing and operating highly available distributed systems. Experience with GPU infrastructure, high-performance computing, ML infrastructure, or large-scale compute platforms. Experience with Kubernetes, cloud infrastructure, Linux, networking, and observability tooling. Excellent debugging, systems design, and operational problem-solving skills. Strong communication skills and experience collaborating across engineering organizations. Equal Opportunity Employment We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. Legal Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations. To notify OpenAI that you believe this job posting is non compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance. We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link. OpenAI Global Applicant Privacy Policy At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.
12/07/2026
Full time
About the Team ChatGPT Engineering builds and operates the compute platform powering one of the world's largest AI products. Every ChatGPT conversation relies on massive GPU clusters serving inference workloads with high reliability, efficiency, and performance. As our GPU fleet continues to grow, we're investing in the infrastructure that operates it. Our team builds the tooling, automation, and intelligent systems that make GPU infrastructure scalable, observable, and increasingly autonomous. We work across production engineering, distributed systems, capacity management, and AI-powered operational tooling to help researchers and product teams move faster while maximizing the efficiency of every GPU. This is a unique opportunity to work on infrastructure at the frontier of AI, where small improvements in fleet efficiency, reliability, and automation have an outsized impact on the development and deployment of AGI. About the Role We're looking for a Software Engineer with deep experience operating large-scale GPU or compute infrastructure. You'll design and build the systems that manage GPU clusters at scale-from fleet health and capacity planning to operational automation and intelligent agents that reduce manual intervention. You'll partner closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and overall compute utilization. This role is ideal for engineers who enjoy solving complex operational challenges, building internal platforms, and working on infrastructure that directly powers frontier AI. In This Role, You Will Design, build, and operate software that manages large-scale GPU infrastructure supporting ChatGPT inference. Build internal platforms, tooling, and AI-powered agents that automate fleet operations and reduce operational overhead. Improve observability, reliability, and operational efficiency across thousands of GPUs. Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response. Identify infrastructure bottlenecks and implement solutions that improve utilization, scalability, and performance. Partner closely with research, platform, networking, and systems teams to continuously improve our compute platform. Help establish engineering best practices around operational excellence, automation, and infrastructure reliability. You Might Thrive in This Role If You Have experience operating large-scale production infrastructure, preferably GPU clusters or other compute-intensive distributed systems. Have a background in Production Engineering, Site Reliability Engineering (SRE), Infrastructure Engineering, or Platform Engineering. Have built software that automates operational workflows rather than relying on manual processes. Have experience with Kubernetes, Linux systems, container orchestration, or distributed infrastructure. Understand infrastructure observability, monitoring, capacity planning, and incident management. Enjoy identifying cross-team pain points and building reusable platforms that improve developer productivity. Are comfortable working across software engineering and systems operations, owning problems end-to-end. Thrive in fast-moving environments with significant technical ambiguity. Qualifications 5+ years of software engineering experience building production infrastructure. Strong programming skills in Go, Python, C++, Rust, or similar systems languages. Experience designing and operating highly available distributed systems. Experience with GPU infrastructure, high-performance computing, ML infrastructure, or large-scale compute platforms. Experience with Kubernetes, cloud infrastructure, Linux, networking, and observability tooling. Excellent debugging, systems design, and operational problem-solving skills. Strong communication skills and experience collaborating across engineering organizations. Equal Opportunity Employment We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. Legal Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations. To notify OpenAI that you believe this job posting is non compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance. We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link. OpenAI Global Applicant Privacy Policy At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.
Enterprise Architect - AI
RadNet, Inc.
Enterprise Architect - Artificial Intelligence AI/ML Infrastructure, Platform & Governance Role Summary World Wide Technology is looking for a deeply technical Enterprise Architect who will own the delivery of AI projects end to end from the silicon and data center design that underpins AI workloads, through the software and MLOps stack, to the governance frameworks that make AI trustworthy and defensible at scale. This is a technical hardware-and-software architect role, not a strategy-only position. The successful candidate operates comfortably across GPU infrastructure, high-performance networking, model training and inference pipelines, and the AI risk/governance disciplines increasingly demanded by regulators and enterprise boards. The Enterprise Architect will lead technical delivery teams for client engagements, acting as the single point of technical accountability from design through to go-live, while mentoring delivery teams and shaping WWT's broader AI point of view. Key Responsibilities Own end-to-end technical delivery of AI/ML engagements: architecture definition, design authority, build oversight, and go-live validation. Host and chair Architecture Review Board (ARB) and Technical Design Authority (TDA) sessions for AI engagements, owning governance gates, decision records, and design sign-off. Architect AI infrastructure spanning GPU/accelerator compute, high-performance interconnects ,parallel/high-throughput storage, and orchestration Design the AI software stack: training and fine-tuning pipelines, distributed training frameworks, inference/serving platforms, MLOps/LLMOps tooling, vector databases, and retrieval-augmented generation (RAG) and agentic architectures. Define and AI governance frameworks covering model risk management, responsible AI, data lineage, bias/fairness testing, explainability, and regulatory alignment (EU AI Act, NIST AI RMF, ISO/IEC 42001). Act as trusted technical advisor to client CTOs, CIOs and Heads of Data/AI on platform strategy, build-vs-buy decisions, and AI operating model design. Lead technical workshops, architecture design sessions, and proof-of-concept builds with cross-functional engineering, data science, and security teams. Serve as the technical escalation point for delivery teams; unblock design and implementation issues under time pressure. Mentor other architects and engineers on AI systems design, uplifting AI depth across the practice. Partner with sales and pre-sales to scope AI solutions, size infrastructure, and validate technical feasibility of proposed architectures. Define automation, orchestration, and observability standards across the AI stack, from GPU cluster provisioning through to model monitoring in production. Architect integration points connecting AI platforms to existing enterprise networks, third-party systems, and external or service-provider-hosted environments (e.g. colocation, managed GPU-as-a-service, external inference endpoints). Track the fast-moving AI landscape - new model architectures, silicon, frameworks and regulation - and translate relevant developments into WWT's delivery methodology and client recommendations. Required Technical Skills & Experience Experience Baseline 10+ years in enterprise architecture, infrastructure engineering, or platform engineering roles. 5+ years focused specifically on AI/ML systems design and delivery, including at least 2 years working with generative AI/LLM workloads. Demonstrated track record leading technical delivery (not just advisory) on enterprise-scale AI or HPC infrastructure programmes. AI Hardware & Data Center Infrastructure GPU/accelerator architectures: NVIDIA / AMD, including multi-node scale-out design. Accelerator interconnects: NVLink, NVSwitch High-performance networking: InfiniBand and RoCEv2 fabric design, 400G/800G Ethernet, rail-optimized topologies for AI clusters. Data center facilities: power density, liquid cooling, and rack-level design considerations specific to AI compute. Storage: parallel and high-throughput file systems (e.g. Everpure, WEKA, VAST, NetApp) sized for training and checkpointing workloads. AI Software, MLOps & Generative AI ML frameworks: PyTorch and TensorFlow at a working, hands-on level. Distributed training: Horovod, DeepSpeed, Megatron-LM, or equivalent multi-node training frameworks. Inference & serving: NVIDIA Triton, vLLM, TensorRT-LLM, or equivalent high-throughput serving platforms. MLOps/LLMOps: Kubeflow, MLflow, and at least one hyperscaler ML platform (SageMaker, Azure ML, or Vertex AI). Generative AI: LLM fine-tuning (LoRA/QLoRA), RAG architecture design, vector databases (Pinecone, Milvus, Weaviate), and agentic frameworks (LangChain, LangGraph, Semantic Kernel). Data pipelines: data lake/lakehouse architectures, ETL/ELT, and data quality/lineage tooling that feed AI systems. Automation, Orchestration & Observability Infrastructure-as-code: Terraform and Ansible for repeatable, automated provisioning of GPU clusters and AI platform environments; GitOps (ArgoCD) for continuous, declarative platform delivery. Pipeline orchestration: Kubeflow Pipelines, Apache Airflow, or Argo Workflows to orchestrate multi-stage training, fine-tuning, and inference pipelines. Cluster & workload scheduling: Slurm, Run:ai, and NVIDIA Base Command Manager for GPU job scheduling; Kubernetes-native GPU scheduling including device plugins and MIG partitioning for multi-tenant clusters. CI/CD/CT for ML: automated model testing, validation gates, and promotion pipelines (continuous training/continuous delivery) that move models safely from experimentation to production. Infrastructure & GPU observability: NVIDIA DCGM, Prometheus/Grafana, and related telemetry stacks for GPU utilization, thermal, and cluster health monitoring. Model & LLM observability: production model performance monitoring, data/concept drift detection, and LLM-specific observability (token usage, latency, cost, hallucination/quality metrics) using tools such as Arize, WhyLabs, or Langfuse. Logging & tracing: centralized logging (ELK/OpenSearch) and distributed tracing (OpenTelemetry) across data, training, and inference pipelines for end-to-end root cause analysis. Integration - AI Stack, Enterprise Networks & Service Provider Environments Platform integration: API-based and event-driven integration of AI platforms with enterprise systems, using REST/gRPC APIs and message/event streaming platforms (e.g. Kafka). Enterprise network integration: designing connectivity between AI/GPU infrastructure and existing campus, data center, and WAN environments, including capacity and latency planning for east west training traffic and north south inference traffic. Hybrid & multi-cloud connectivity: integrating on-premises AI platforms with cloud AI services via dedicated interconnects (Direct Connect, ExpressRoute) and multi-cloud/hybrid connectivity patterns for distributed training or burst inference. Service provider & third party integration: experience architecting connections into external or service provider hosted environments - colocation interconnects, managed GPU-as-a-service offerings, and third party/external inference endpoints - including the commercial and technical boundary considerations involved. Secure exposure of AI services: working knowledge of API gateways, service mesh, and mutual TLS as applied to exposing or consuming AI services safely across organizational and network boundaries. Cross-functional design: proven ability to partner directly with network and security architects to define end to end integration architecture spanning AI platforms, enterprise networks, and external/customer environments. Working knowledge of model risk management frameworks and responsible AI principles (fairness, explainability, human oversight). Familiarity with data privacy regulation (GDPR, CCPA) as applied to AI training and inference data. Working knowledge of emerging AI specific regulation and standards: EU AI Act, NIST AI Risk Management Framework, ISO/IEC 42001. Experience establishing model documentation, audit trail, and approval gate processes for production AI systems. Security & Cloud AI specific security fundamentals: model security, prompt injection defenses, supply chain security for open source/open weight models. Solutions architect level expertise in at least one hyperscaler (AWS, Azure, or GCP), including their native AI/ML services. Ability to design for hybrid on premises/cloud AI deployments, including data residency and sovereignty constraints. Architecture Governance & Design Authority Proven experience hosting and chairing formal Architecture Review Board (ARB) and Technical Design Authority (TDA) forums, including agenda ownership, decision logging, and stakeholder facilitation. Ability to define and operate governance gates across the engagement lifecycle: design authority sign off, change control, and exception/waiver management for AI platform decisions. Experience producing and maintaining architecture decision records (ADRs), design standards, and reference architectures that are actively enforced through ARB/TDA governance. . click apply for full job details
12/07/2026
Full time
Enterprise Architect - Artificial Intelligence AI/ML Infrastructure, Platform & Governance Role Summary World Wide Technology is looking for a deeply technical Enterprise Architect who will own the delivery of AI projects end to end from the silicon and data center design that underpins AI workloads, through the software and MLOps stack, to the governance frameworks that make AI trustworthy and defensible at scale. This is a technical hardware-and-software architect role, not a strategy-only position. The successful candidate operates comfortably across GPU infrastructure, high-performance networking, model training and inference pipelines, and the AI risk/governance disciplines increasingly demanded by regulators and enterprise boards. The Enterprise Architect will lead technical delivery teams for client engagements, acting as the single point of technical accountability from design through to go-live, while mentoring delivery teams and shaping WWT's broader AI point of view. Key Responsibilities Own end-to-end technical delivery of AI/ML engagements: architecture definition, design authority, build oversight, and go-live validation. Host and chair Architecture Review Board (ARB) and Technical Design Authority (TDA) sessions for AI engagements, owning governance gates, decision records, and design sign-off. Architect AI infrastructure spanning GPU/accelerator compute, high-performance interconnects ,parallel/high-throughput storage, and orchestration Design the AI software stack: training and fine-tuning pipelines, distributed training frameworks, inference/serving platforms, MLOps/LLMOps tooling, vector databases, and retrieval-augmented generation (RAG) and agentic architectures. Define and AI governance frameworks covering model risk management, responsible AI, data lineage, bias/fairness testing, explainability, and regulatory alignment (EU AI Act, NIST AI RMF, ISO/IEC 42001). Act as trusted technical advisor to client CTOs, CIOs and Heads of Data/AI on platform strategy, build-vs-buy decisions, and AI operating model design. Lead technical workshops, architecture design sessions, and proof-of-concept builds with cross-functional engineering, data science, and security teams. Serve as the technical escalation point for delivery teams; unblock design and implementation issues under time pressure. Mentor other architects and engineers on AI systems design, uplifting AI depth across the practice. Partner with sales and pre-sales to scope AI solutions, size infrastructure, and validate technical feasibility of proposed architectures. Define automation, orchestration, and observability standards across the AI stack, from GPU cluster provisioning through to model monitoring in production. Architect integration points connecting AI platforms to existing enterprise networks, third-party systems, and external or service-provider-hosted environments (e.g. colocation, managed GPU-as-a-service, external inference endpoints). Track the fast-moving AI landscape - new model architectures, silicon, frameworks and regulation - and translate relevant developments into WWT's delivery methodology and client recommendations. Required Technical Skills & Experience Experience Baseline 10+ years in enterprise architecture, infrastructure engineering, or platform engineering roles. 5+ years focused specifically on AI/ML systems design and delivery, including at least 2 years working with generative AI/LLM workloads. Demonstrated track record leading technical delivery (not just advisory) on enterprise-scale AI or HPC infrastructure programmes. AI Hardware & Data Center Infrastructure GPU/accelerator architectures: NVIDIA / AMD, including multi-node scale-out design. Accelerator interconnects: NVLink, NVSwitch High-performance networking: InfiniBand and RoCEv2 fabric design, 400G/800G Ethernet, rail-optimized topologies for AI clusters. Data center facilities: power density, liquid cooling, and rack-level design considerations specific to AI compute. Storage: parallel and high-throughput file systems (e.g. Everpure, WEKA, VAST, NetApp) sized for training and checkpointing workloads. AI Software, MLOps & Generative AI ML frameworks: PyTorch and TensorFlow at a working, hands-on level. Distributed training: Horovod, DeepSpeed, Megatron-LM, or equivalent multi-node training frameworks. Inference & serving: NVIDIA Triton, vLLM, TensorRT-LLM, or equivalent high-throughput serving platforms. MLOps/LLMOps: Kubeflow, MLflow, and at least one hyperscaler ML platform (SageMaker, Azure ML, or Vertex AI). Generative AI: LLM fine-tuning (LoRA/QLoRA), RAG architecture design, vector databases (Pinecone, Milvus, Weaviate), and agentic frameworks (LangChain, LangGraph, Semantic Kernel). Data pipelines: data lake/lakehouse architectures, ETL/ELT, and data quality/lineage tooling that feed AI systems. Automation, Orchestration & Observability Infrastructure-as-code: Terraform and Ansible for repeatable, automated provisioning of GPU clusters and AI platform environments; GitOps (ArgoCD) for continuous, declarative platform delivery. Pipeline orchestration: Kubeflow Pipelines, Apache Airflow, or Argo Workflows to orchestrate multi-stage training, fine-tuning, and inference pipelines. Cluster & workload scheduling: Slurm, Run:ai, and NVIDIA Base Command Manager for GPU job scheduling; Kubernetes-native GPU scheduling including device plugins and MIG partitioning for multi-tenant clusters. CI/CD/CT for ML: automated model testing, validation gates, and promotion pipelines (continuous training/continuous delivery) that move models safely from experimentation to production. Infrastructure & GPU observability: NVIDIA DCGM, Prometheus/Grafana, and related telemetry stacks for GPU utilization, thermal, and cluster health monitoring. Model & LLM observability: production model performance monitoring, data/concept drift detection, and LLM-specific observability (token usage, latency, cost, hallucination/quality metrics) using tools such as Arize, WhyLabs, or Langfuse. Logging & tracing: centralized logging (ELK/OpenSearch) and distributed tracing (OpenTelemetry) across data, training, and inference pipelines for end-to-end root cause analysis. Integration - AI Stack, Enterprise Networks & Service Provider Environments Platform integration: API-based and event-driven integration of AI platforms with enterprise systems, using REST/gRPC APIs and message/event streaming platforms (e.g. Kafka). Enterprise network integration: designing connectivity between AI/GPU infrastructure and existing campus, data center, and WAN environments, including capacity and latency planning for east west training traffic and north south inference traffic. Hybrid & multi-cloud connectivity: integrating on-premises AI platforms with cloud AI services via dedicated interconnects (Direct Connect, ExpressRoute) and multi-cloud/hybrid connectivity patterns for distributed training or burst inference. Service provider & third party integration: experience architecting connections into external or service provider hosted environments - colocation interconnects, managed GPU-as-a-service offerings, and third party/external inference endpoints - including the commercial and technical boundary considerations involved. Secure exposure of AI services: working knowledge of API gateways, service mesh, and mutual TLS as applied to exposing or consuming AI services safely across organizational and network boundaries. Cross-functional design: proven ability to partner directly with network and security architects to define end to end integration architecture spanning AI platforms, enterprise networks, and external/customer environments. Working knowledge of model risk management frameworks and responsible AI principles (fairness, explainability, human oversight). Familiarity with data privacy regulation (GDPR, CCPA) as applied to AI training and inference data. Working knowledge of emerging AI specific regulation and standards: EU AI Act, NIST AI Risk Management Framework, ISO/IEC 42001. Experience establishing model documentation, audit trail, and approval gate processes for production AI systems. Security & Cloud AI specific security fundamentals: model security, prompt injection defenses, supply chain security for open source/open weight models. Solutions architect level expertise in at least one hyperscaler (AWS, Azure, or GCP), including their native AI/ML services. Ability to design for hybrid on premises/cloud AI deployments, including data residency and sovereignty constraints. Architecture Governance & Design Authority Proven experience hosting and chairing formal Architecture Review Board (ARB) and Technical Design Authority (TDA) forums, including agenda ownership, decision logging, and stakeholder facilitation. Ability to define and operate governance gates across the engagement lifecycle: design authority sign off, change control, and exception/waiver management for AI platform decisions. Experience producing and maintaining architecture decision records (ADRs), design standards, and reference architectures that are actively enforced through ARB/TDA governance. . click apply for full job details
Distributed Scheduling & Workload Orchestration Engineer
Viridiengroup Crawley, Sussex
A technology and data company in Crawley seeks a Software Developer focusing on distributed scheduling and workload orchestration. This position requires strong software development skills and experience with backend services and job scheduling systems. The developer will design systems for workload orchestration using tools such as Slurm, PostgreSQL, and microservices. The company offers competitive salary, bonuses, and flexible benefits, creating a supportive and innovative work environment.
12/07/2026
Full time
A technology and data company in Crawley seeks a Software Developer focusing on distributed scheduling and workload orchestration. This position requires strong software development skills and experience with backend services and job scheduling systems. The developer will design systems for workload orchestration using tools such as Slurm, PostgreSQL, and microservices. The company offers competitive salary, bonuses, and flexible benefits, creating a supportive and innovative work environment.
DevOps Engineer
Diffractive Labs
What We're Looking For We are seeking a DevOps Engineer to build and own the infrastructure that underpins our AI driven materials discovery platform. You'll work directly with world renowned ML researchers and software engineers to accelerate real scientific breakthroughs by making model training, experimentation, and deployment fast, reliable, and reproducible. This is a foundational hire. You'll set the patterns others build on. You will be joining a small, highly ambitious team of world renowned engineers, AI researchers, and materials scientists. We move fast and value people who are energised by that. What You'll Do Design, provision, and manage cloud infrastructure (AWS/GCP) using infrastructure as code; Terraform, Pulumi, or equivalent. Own GPU compute environments for model training and inference, including cluster configuration, job scheduling, and cost optimisation. Build and maintain CI/CD pipelines that support rapid model iteration, automated testing, and safe deployments. Support ML workflow orchestration; experiment tracking, training run management, and data pipeline reliability. Ensure reproducibility across research and production environments through containerisation and rigorous environment management. Define monitoring, alerting, and incident response processes so the team can move fast without things silently breaking. Implement security best practices: secrets management, IAM, network segmentation, vulnerability scanning. Build internal tooling and documentation that lets researchers self serve infrastructure without waiting on you. Skills & Qualifications 4+ years in a DevOps, Platform Engineering, or SRE role. Strong proficiency with at least one major cloud provider and its core services (compute, storage, networking, IAM). Hands on experience with infrastructure as code and container orchestration (Kubernetes or equivalent). Solid CI/CD pipeline experience, GitHub Actions, GitLab CI, or similar. Proficient in Python and Bash; comfortable reading and writing code across a polyglot stack. Deep Linux systems knowledge and strong networking fundamentals. A bias for building things properly the first time, even under early stage constraints. Nice to Have Experience with GPU cluster management and ML training workloads (NVIDIA, CUDA, distributed training). Familiarity with MLOps tooling: Experiment tracking (MLflow, Weights & Biases). Workflow orchestration (Airflow, Prefect, Argo). Data versioning (DVC). Background in scientific computing or HPC environments. Prior experience at a deep tech or computational science company. Why Join Us Work directly on infrastructure that enables AI to make real scientific discoveries. Shape how we build from day one, no legacy systems, no inherited mess. Collaborate with world class researchers across materials science and machine learning. Diffractive is building the AI Material Scientist that autonomously learns from real world experimentation to push the boundaries of scientific discovery. We're early, moving fast, and working on problems that genuinely matter. We are a London based company with a flexible approach to how and where you work. We offer competitive salary, generous equity, and benefits. You'll have a real stake in what you build and in the company's overall success. Equal Opportunity Diffractive is an equal opportunities employer. We are committed to creating an inclusive environment for all employees and welcome applications from people of all backgrounds, experiences, and identities. If you require any adjustments or accommodations at any point during the interview process please let us know - we will be happy to help.
11/07/2026
Full time
What We're Looking For We are seeking a DevOps Engineer to build and own the infrastructure that underpins our AI driven materials discovery platform. You'll work directly with world renowned ML researchers and software engineers to accelerate real scientific breakthroughs by making model training, experimentation, and deployment fast, reliable, and reproducible. This is a foundational hire. You'll set the patterns others build on. You will be joining a small, highly ambitious team of world renowned engineers, AI researchers, and materials scientists. We move fast and value people who are energised by that. What You'll Do Design, provision, and manage cloud infrastructure (AWS/GCP) using infrastructure as code; Terraform, Pulumi, or equivalent. Own GPU compute environments for model training and inference, including cluster configuration, job scheduling, and cost optimisation. Build and maintain CI/CD pipelines that support rapid model iteration, automated testing, and safe deployments. Support ML workflow orchestration; experiment tracking, training run management, and data pipeline reliability. Ensure reproducibility across research and production environments through containerisation and rigorous environment management. Define monitoring, alerting, and incident response processes so the team can move fast without things silently breaking. Implement security best practices: secrets management, IAM, network segmentation, vulnerability scanning. Build internal tooling and documentation that lets researchers self serve infrastructure without waiting on you. Skills & Qualifications 4+ years in a DevOps, Platform Engineering, or SRE role. Strong proficiency with at least one major cloud provider and its core services (compute, storage, networking, IAM). Hands on experience with infrastructure as code and container orchestration (Kubernetes or equivalent). Solid CI/CD pipeline experience, GitHub Actions, GitLab CI, or similar. Proficient in Python and Bash; comfortable reading and writing code across a polyglot stack. Deep Linux systems knowledge and strong networking fundamentals. A bias for building things properly the first time, even under early stage constraints. Nice to Have Experience with GPU cluster management and ML training workloads (NVIDIA, CUDA, distributed training). Familiarity with MLOps tooling: Experiment tracking (MLflow, Weights & Biases). Workflow orchestration (Airflow, Prefect, Argo). Data versioning (DVC). Background in scientific computing or HPC environments. Prior experience at a deep tech or computational science company. Why Join Us Work directly on infrastructure that enables AI to make real scientific discoveries. Shape how we build from day one, no legacy systems, no inherited mess. Collaborate with world class researchers across materials science and machine learning. Diffractive is building the AI Material Scientist that autonomously learns from real world experimentation to push the boundaries of scientific discovery. We're early, moving fast, and working on problems that genuinely matter. We are a London based company with a flexible approach to how and where you work. We offer competitive salary, generous equity, and benefits. You'll have a real stake in what you build and in the company's overall success. Equal Opportunity Diffractive is an equal opportunities employer. We are committed to creating an inclusive environment for all employees and welcome applications from people of all backgrounds, experiences, and identities. If you require any adjustments or accommodations at any point during the interview process please let us know - we will be happy to help.
Member of Technical Staff (AI Inference Engineer)
CVFine by Instrovate Technologies
# Member of Technical Staff (AI Inference Engineer)PerplexityVia company siteLondon, UK, United Kingdom 56 days ago 0 interestedAIPython Job DescriptionWe are looking for an AI Inference Engineer to join our growing team. We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL. Responsibilities: New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway. GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow. Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic. Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernels interleaving. Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents. Who we're looking for: Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar). Any other deep systems programming experience is a plus. You understand modern LLM architectures and are able to bring them up reliably in a production environment. You've built and operated production distributed systems under real load - ideally performance-critical ones. Comfortable working across languages and layers: Rust for the serving runtime, Python for model code, CUDA/CuteDSL for kernels. You own problems end-to-end. You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on Friday. Self-directed. You do well in fast-moving environments where the path forward isn't laid out for you. Nice-to-have: ML compilers and framework internals: PyTorch internals, torch.compile, custom operators. Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism. Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving. Profiling and debugging tools: Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis. Container orchestration: Kubernetes, GPU scheduling, autoscaling inference workloads. Qualifications: 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems. Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow). Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores). Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation).Final offer amounts are determined by multiple factors including experience and expertise. Equity: In addition to the base salary, equity may be part of the total compensation package.
10/07/2026
Full time
# Member of Technical Staff (AI Inference Engineer)PerplexityVia company siteLondon, UK, United Kingdom 56 days ago 0 interestedAIPython Job DescriptionWe are looking for an AI Inference Engineer to join our growing team. We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL. Responsibilities: New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway. GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow. Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic. Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernels interleaving. Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents. Who we're looking for: Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar). Any other deep systems programming experience is a plus. You understand modern LLM architectures and are able to bring them up reliably in a production environment. You've built and operated production distributed systems under real load - ideally performance-critical ones. Comfortable working across languages and layers: Rust for the serving runtime, Python for model code, CUDA/CuteDSL for kernels. You own problems end-to-end. You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on Friday. Self-directed. You do well in fast-moving environments where the path forward isn't laid out for you. Nice-to-have: ML compilers and framework internals: PyTorch internals, torch.compile, custom operators. Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism. Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving. Profiling and debugging tools: Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis. Container orchestration: Kubernetes, GPU scheduling, autoscaling inference workloads. Qualifications: 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems. Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow). Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores). Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation).Final offer amounts are determined by multiple factors including experience and expertise. Equity: In addition to the base salary, equity may be part of the total compensation package.
ML Infrastructure Engineer
White Circle
TLDR: We are looking for an ML Infrastructure Engineer to build the systems behind our LLM post training, RL, evaluation, inference, and agentic development workflows. You will work close to researchers, GPUs, training loops, data control systems, evals, inference stacks, and the infrastructure decisions that directly affect model learning and product quality. About us White Circle is an AI Safety company building the safety, reliability, and optimization layer for AI systems. At the core of our platform are policies - simple natural language rules that define what an AI model should and shouldn't do. We automatically test, enforce, and continuously improve these policies at scale. We've raised $11M from top funds, founders, and senior leaders at OpenAI, Anthropic, HuggingFace, Mistral, DeepMind, Datadog, Sentry, and others We process over 100M+ API calls every month We fine tune and train our own LLMs so they run faster and cheaper than any open or proprietary model We're a small, highly focused team. If you want to work deeply on hard problems, see your work ship to production quickly, and influence how AI safety is actually built - you're the one we need. You will Build robust, flexible, and scalable RL and post training pipelines, including smoke tuning runs for quality testing and approach ablations Design data control systems that govern what the model sees, when it sees it, and how training data flows through rollouts, replay, filtering, evaluation, and policy updates Tune training and inference end to end for high throughput across the systems that matter: networking, memory, compute scheduling, data loading, storage, checkpointing, and I/O Investigate how infrastructure choices affect learning dynamics, eval quality, model behavior, and training stability - staying close to the state of the art in LLMs, RL, and post training Build infrastructure for model iteration: experiment runs, artifacts, evals, dashboards, failure inspection, reproducibility, and cost visibility Work on inference infrastructure where it affects post training and evaluation loops Build and improve agentic development environments: coding agent harnesses, browser/tool integrations, terminal/runtime sandboxes, repo aware workflows, and multi agent orchestration Work closely with the team: plan future steps, discuss tradeoffs, share context early, and stay in touch while building You'll fit right in if you Have designed, built, or maintained distributed RL/post training systems at scale and are fluent in their moving parts: rollouts, replay buffers, reward signals, data filtering, policy updates, evaluation loops, and failure analysis Are familiar with deep learning frameworks such as PyTorch or JAX Are proficient in Python, including concurrency, asynchronous programming, multiprocessing, and performance optimization Can debug distributed GPU workloads across CUDA runtime, container runtime, driver versions, NCCL or equivalent communication layers, networking, storage, scheduling, and checkpointing Have experience with profiling tools across the stack, for example py spy, PyTorch profiler, Nsight, perf, tracing, metrics, logs, or custom instrumentation Have experience with inference stacks such as vLLM, SGLang, TensorRT LLM, Dynamo, or custom serving infrastructure Can reason from system metrics back to model behavior: when latency, queueing, sampling, data order, rollout throughput, or infrastructure failures affect learning Have a strong ownership mindset: you can take an ambiguous infrastructure problem, make it concrete, ship a working system, and improve it from real feedback A big plus A public builder footprint: open source contributions to RL, distributed ML, LLM training, inference, eval, or agent infrastructure - repos, PRs, benchmarks, papers with code, technical posts - and a good technical X/Twitter presence with live building, debugging threads, and useful interaction with strong builders Experience in a high bar AI infra, research, or model environment such as xAI/Grok, Qwen, ByteDance AI infra/research, Prime Intellect, or similar teams Custom training framework support or ownership: distributed training, fine tuning pipelines, trainers, schedulers, checkpointing, data loaders, model/eval integration, or performance tooling Serious use of Claude Code, Codex, Kimi Code, Pi Agent, Droid, or similar agentic coding systems as a development surface Experience with GPU clusters on Kubernetes, Slurm, Ray, custom schedulers, or cloud GPU orchestration NCCL, UCX, NVSHMEM, RDMA, InfiniBand, RoCE, or EFA Rust, C++, CUDA, Go, or systems level performance work Why White Circle You will be able to propose and run your own experiments and research ideas on modern ML infrastructure with very little friction You will work on current ML infra problems, close to research, product needs, and real model iteration - not maintaining legacy systems for the sake of keeping them alive You will have an unusually high contribution level for the size of the team; your systems decisions can change how quickly we train, evaluate, ship, and improve models You will have room to dig into areas of your own interest, as long as they help the company build better, faster, safer AI systems Paid time off in line with your local regulations, no matter where you work from. Work from Paris (hybrid) with a relocation package available, or work from London (note: we are currently unable to provide relocation support and medical insurance for London based roles). Comprehensive medical insurance for our France based team. All the hardware, tools, and services you need. Covered subscriptions for AI agents and IDEs. Team off sites twice a year: we've recently been to the Alps and to Saint Tropez.
10/07/2026
Full time
TLDR: We are looking for an ML Infrastructure Engineer to build the systems behind our LLM post training, RL, evaluation, inference, and agentic development workflows. You will work close to researchers, GPUs, training loops, data control systems, evals, inference stacks, and the infrastructure decisions that directly affect model learning and product quality. About us White Circle is an AI Safety company building the safety, reliability, and optimization layer for AI systems. At the core of our platform are policies - simple natural language rules that define what an AI model should and shouldn't do. We automatically test, enforce, and continuously improve these policies at scale. We've raised $11M from top funds, founders, and senior leaders at OpenAI, Anthropic, HuggingFace, Mistral, DeepMind, Datadog, Sentry, and others We process over 100M+ API calls every month We fine tune and train our own LLMs so they run faster and cheaper than any open or proprietary model We're a small, highly focused team. If you want to work deeply on hard problems, see your work ship to production quickly, and influence how AI safety is actually built - you're the one we need. You will Build robust, flexible, and scalable RL and post training pipelines, including smoke tuning runs for quality testing and approach ablations Design data control systems that govern what the model sees, when it sees it, and how training data flows through rollouts, replay, filtering, evaluation, and policy updates Tune training and inference end to end for high throughput across the systems that matter: networking, memory, compute scheduling, data loading, storage, checkpointing, and I/O Investigate how infrastructure choices affect learning dynamics, eval quality, model behavior, and training stability - staying close to the state of the art in LLMs, RL, and post training Build infrastructure for model iteration: experiment runs, artifacts, evals, dashboards, failure inspection, reproducibility, and cost visibility Work on inference infrastructure where it affects post training and evaluation loops Build and improve agentic development environments: coding agent harnesses, browser/tool integrations, terminal/runtime sandboxes, repo aware workflows, and multi agent orchestration Work closely with the team: plan future steps, discuss tradeoffs, share context early, and stay in touch while building You'll fit right in if you Have designed, built, or maintained distributed RL/post training systems at scale and are fluent in their moving parts: rollouts, replay buffers, reward signals, data filtering, policy updates, evaluation loops, and failure analysis Are familiar with deep learning frameworks such as PyTorch or JAX Are proficient in Python, including concurrency, asynchronous programming, multiprocessing, and performance optimization Can debug distributed GPU workloads across CUDA runtime, container runtime, driver versions, NCCL or equivalent communication layers, networking, storage, scheduling, and checkpointing Have experience with profiling tools across the stack, for example py spy, PyTorch profiler, Nsight, perf, tracing, metrics, logs, or custom instrumentation Have experience with inference stacks such as vLLM, SGLang, TensorRT LLM, Dynamo, or custom serving infrastructure Can reason from system metrics back to model behavior: when latency, queueing, sampling, data order, rollout throughput, or infrastructure failures affect learning Have a strong ownership mindset: you can take an ambiguous infrastructure problem, make it concrete, ship a working system, and improve it from real feedback A big plus A public builder footprint: open source contributions to RL, distributed ML, LLM training, inference, eval, or agent infrastructure - repos, PRs, benchmarks, papers with code, technical posts - and a good technical X/Twitter presence with live building, debugging threads, and useful interaction with strong builders Experience in a high bar AI infra, research, or model environment such as xAI/Grok, Qwen, ByteDance AI infra/research, Prime Intellect, or similar teams Custom training framework support or ownership: distributed training, fine tuning pipelines, trainers, schedulers, checkpointing, data loaders, model/eval integration, or performance tooling Serious use of Claude Code, Codex, Kimi Code, Pi Agent, Droid, or similar agentic coding systems as a development surface Experience with GPU clusters on Kubernetes, Slurm, Ray, custom schedulers, or cloud GPU orchestration NCCL, UCX, NVSHMEM, RDMA, InfiniBand, RoCE, or EFA Rust, C++, CUDA, Go, or systems level performance work Why White Circle You will be able to propose and run your own experiments and research ideas on modern ML infrastructure with very little friction You will work on current ML infra problems, close to research, product needs, and real model iteration - not maintaining legacy systems for the sake of keeping them alive You will have an unusually high contribution level for the size of the team; your systems decisions can change how quickly we train, evaluate, ship, and improve models You will have room to dig into areas of your own interest, as long as they help the company build better, faster, safer AI systems Paid time off in line with your local regulations, no matter where you work from. Work from Paris (hybrid) with a relocation package available, or work from London (note: we are currently unable to provide relocation support and medical insurance for London based roles). Comprehensive medical insurance for our France based team. All the hardware, tools, and services you need. Covered subscriptions for AI agents and IDEs. Team off sites twice a year: we've recently been to the Alps and to Saint Tropez.
Member of Technical Staff (AI Infrastructure Engineer)
Aimling
Overview We are looking for an AI Infra engineer to join our growing team. We work with Kubernetes, Slurm, Python, C++, PyTorch, and primarily on AWS. As an AI Infrastructure Engineer, you will be partnering closely with our Inference and Research teams to build, deploy, and optimize our large-scale AI training and inference clusters. Responsibilities Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads Manage and optimize Slurm-based HPC environments for distributed training of large language models Develop robust APIs and orchestration systems for both training pipelines and inference services Implement resource scheduling and job management systems across heterogeneous compute environments Benchmark system performance, diagnose bottlenecks, and implement improvements across both training and inference infrastructure Build monitoring, alerting, and observability solutions tailored to ML workloads running on Kubernetes and Slurm Respond swiftly to system outages and collaborate across teams to maintain high uptime for critical training runs and inference services Optimize cluster utilization and implement autoscaling strategies for dynamic workload demands Qualifications Strong expertise in Kubernetes administration, including custom resource definitions, operators, and cluster management Hands on experience with Slurm workload management, including job scheduling, resource allocation, and cluster optimization Experience with deploying and managing distributed training systems at scale Deep understanding of container orchestration and distributed systems architecture High level familiarity with LLM architecture and training processes (Multi Head Attention, Multi/Grouped Query, distributed training strategies) Experience managing GPU clusters and optimizing compute resource utilization Required Skills Expert level Kubernetes administration and YAML configuration management Proficiency with Slurm job scheduling, resource management, and cluster configuration Python and C++ programming with focus on systems and infrastructure automation Hands on experience with ML frameworks such as PyTorch in distributed training contexts Strong understanding of networking, storage, and compute resource management for ML workloads Experience developing APIs and managing distributed systems for both batch and real time workloads Solid debugging and monitoring skills with expertise in observability tools for containerized environments Preferred Skills Experience with Kubernetes operators and custom controllers for ML workloads Advanced Slurm administration including multi cluster federation and advanced scheduling policies Familiarity with GPU cluster management and CUDA optimization Experience with other ML frameworks like TensorFlow or distributed training libraries Background in HPC environments, parallel computing, and high performance networking Knowledge of infrastructure as code (Terraform, Ansible) and GitOps practices Experience with container registries, image optimization, and multi stage builds for ML workloads Required Experience Demonstrated experience managing large scale Kubernetes deployments in production environments Proven track record with Slurm cluster administration and HPC workload management Previous roles in SRE, DevOps, or Platform Engineering with focus on ML infrastructure Experience supporting both long running training jobs and high availability inference services Ideally, 3-5 years of relevant experience in ML systems deployment with specific focus on cluster orchestration and resource management
06/07/2026
Full time
Overview We are looking for an AI Infra engineer to join our growing team. We work with Kubernetes, Slurm, Python, C++, PyTorch, and primarily on AWS. As an AI Infrastructure Engineer, you will be partnering closely with our Inference and Research teams to build, deploy, and optimize our large-scale AI training and inference clusters. Responsibilities Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads Manage and optimize Slurm-based HPC environments for distributed training of large language models Develop robust APIs and orchestration systems for both training pipelines and inference services Implement resource scheduling and job management systems across heterogeneous compute environments Benchmark system performance, diagnose bottlenecks, and implement improvements across both training and inference infrastructure Build monitoring, alerting, and observability solutions tailored to ML workloads running on Kubernetes and Slurm Respond swiftly to system outages and collaborate across teams to maintain high uptime for critical training runs and inference services Optimize cluster utilization and implement autoscaling strategies for dynamic workload demands Qualifications Strong expertise in Kubernetes administration, including custom resource definitions, operators, and cluster management Hands on experience with Slurm workload management, including job scheduling, resource allocation, and cluster optimization Experience with deploying and managing distributed training systems at scale Deep understanding of container orchestration and distributed systems architecture High level familiarity with LLM architecture and training processes (Multi Head Attention, Multi/Grouped Query, distributed training strategies) Experience managing GPU clusters and optimizing compute resource utilization Required Skills Expert level Kubernetes administration and YAML configuration management Proficiency with Slurm job scheduling, resource management, and cluster configuration Python and C++ programming with focus on systems and infrastructure automation Hands on experience with ML frameworks such as PyTorch in distributed training contexts Strong understanding of networking, storage, and compute resource management for ML workloads Experience developing APIs and managing distributed systems for both batch and real time workloads Solid debugging and monitoring skills with expertise in observability tools for containerized environments Preferred Skills Experience with Kubernetes operators and custom controllers for ML workloads Advanced Slurm administration including multi cluster federation and advanced scheduling policies Familiarity with GPU cluster management and CUDA optimization Experience with other ML frameworks like TensorFlow or distributed training libraries Background in HPC environments, parallel computing, and high performance networking Knowledge of infrastructure as code (Terraform, Ansible) and GitOps practices Experience with container registries, image optimization, and multi stage builds for ML workloads Required Experience Demonstrated experience managing large scale Kubernetes deployments in production environments Proven track record with Slurm cluster administration and HPC workload management Previous roles in SRE, DevOps, or Platform Engineering with focus on ML infrastructure Experience supporting both long running training jobs and high availability inference services Ideally, 3-5 years of relevant experience in ML systems deployment with specific focus on cluster orchestration and resource management

Modal Window

  • Home
  • Contact
  • About Us
  • FAQs
  • Terms & Conditions
  • Privacy
  • Employer
  • Post a Job
  • Search Resumes
  • Sign in
  • Job Seeker
  • Find Jobs
  • Create Resume
  • Sign in
  • IT blog
  • Facebook
  • Twitter
  • LinkedIn
  • Youtube
© 2008-2026 IT Job Board