it job board logo
  • Home
  • Find IT Jobs
  • Register CV
  • Career Advice
  • Contact us
  • Employers
    • Register as Employer
    • Pricing Plans
  • Recruiting? Post a job
  • Sign in
  • Sign up
  • Home
  • Find IT Jobs
  • Register CV
  • Career Advice
  • Contact us
  • Employers
    • Register as Employer
    • Pricing Plans
Sorry, that job is no longer available. Here are some results that may be similar to the job you were looking for.

11 jobs found

Email me jobs like this
Refine Search
Current Search
software developer distributed scheduling workload orchestration
Software Engineer, GPU Infrastructure- ChatGPT Engineering
Doist
About the Team: ChatGPT Engineering builds and operates the compute platform powering one of the world's largest AI products. Every ChatGPT conversation relies on massive GPU clusters serving inference workloads with high reliability, efficiency, and performance. As our GPU fleet continues to grow, we're investing in the infrastructure that operates it. Our team builds the tooling, automation, and intelligent systems that make GPU infrastructure scalable, observable, and increasingly autonomous. We work across production engineering, distributed systems, capacity management, and AI powered operational tooling to help researchers and product teams move faster while maximizing the efficiency of every GPU. About the Role We're looking for a Software Engineer with deep experience operating large scale GPU or compute infrastructure. You'll design and build the systems that manage GPU clusters at scale-from fleet health and capacity planning to operational automation and intelligent agents that reduce manual intervention. You'll partner closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and overall compute utilization. This role is ideal for engineers who enjoy solving complex operational challenges, building internal platforms, and working on infrastructure that directly powers frontier AI. In This Role, You Will Design, build, and operate software that manages large scale GPU infrastructure supporting ChatGPT inference. Build internal platforms, tooling, and AI powered agents that automate fleet operations and reduce operational overhead. Improve observability, reliability, and operational efficiency across thousands of GPUs. Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response. Identify infrastructure bottlenecks and implement solutions that improve utilization, scalability, and performance. Partner closely with research, platform, networking, and systems teams to continuously improve our compute platform. Help establish engineering best practices around operational excellence, automation, and infrastructure reliability. You Might Thrive in This Role Have experience operating large scale production infrastructure, preferably GPU clusters or other compute intensive distributed systems. Have a background in Production Engineering, Site Reliability Engineering (SRE), Infrastructure Engineering, or Platform Engineering. Have built software that automates operational workflows rather than relying on manual processes. Have experience with Kubernetes, Linux systems, container orchestration, or distributed infrastructure. Understand infrastructure observability, monitoring, capacity planning, and incident management. Enjoy identifying cross team pain points and building reusable platforms that improve developer productivity. Are comfortable working across software engineering and systems operations, owning problems end to end. Thrive in fast moving environments with significant technical ambiguity. Qualifications 5+ years of software engineering experience building production infrastructure. Strong programming skills in Go, Python, C++, Rust, or similar systems languages. Experience designing and operating highly available distributed systems. Experience with GPU infrastructure, high performance computing, ML infrastructure, or large scale compute platforms. Experience with Kubernetes, cloud infrastructure, Linux, networking, and observability tooling. Excellent debugging, systems design, and operational problem solving skills. Strong communication skills and experience collaborating across engineering organizations. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
23/07/2026
Full time
About the Team: ChatGPT Engineering builds and operates the compute platform powering one of the world's largest AI products. Every ChatGPT conversation relies on massive GPU clusters serving inference workloads with high reliability, efficiency, and performance. As our GPU fleet continues to grow, we're investing in the infrastructure that operates it. Our team builds the tooling, automation, and intelligent systems that make GPU infrastructure scalable, observable, and increasingly autonomous. We work across production engineering, distributed systems, capacity management, and AI powered operational tooling to help researchers and product teams move faster while maximizing the efficiency of every GPU. About the Role We're looking for a Software Engineer with deep experience operating large scale GPU or compute infrastructure. You'll design and build the systems that manage GPU clusters at scale-from fleet health and capacity planning to operational automation and intelligent agents that reduce manual intervention. You'll partner closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and overall compute utilization. This role is ideal for engineers who enjoy solving complex operational challenges, building internal platforms, and working on infrastructure that directly powers frontier AI. In This Role, You Will Design, build, and operate software that manages large scale GPU infrastructure supporting ChatGPT inference. Build internal platforms, tooling, and AI powered agents that automate fleet operations and reduce operational overhead. Improve observability, reliability, and operational efficiency across thousands of GPUs. Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response. Identify infrastructure bottlenecks and implement solutions that improve utilization, scalability, and performance. Partner closely with research, platform, networking, and systems teams to continuously improve our compute platform. Help establish engineering best practices around operational excellence, automation, and infrastructure reliability. You Might Thrive in This Role Have experience operating large scale production infrastructure, preferably GPU clusters or other compute intensive distributed systems. Have a background in Production Engineering, Site Reliability Engineering (SRE), Infrastructure Engineering, or Platform Engineering. Have built software that automates operational workflows rather than relying on manual processes. Have experience with Kubernetes, Linux systems, container orchestration, or distributed infrastructure. Understand infrastructure observability, monitoring, capacity planning, and incident management. Enjoy identifying cross team pain points and building reusable platforms that improve developer productivity. Are comfortable working across software engineering and systems operations, owning problems end to end. Thrive in fast moving environments with significant technical ambiguity. Qualifications 5+ years of software engineering experience building production infrastructure. Strong programming skills in Go, Python, C++, Rust, or similar systems languages. Experience designing and operating highly available distributed systems. Experience with GPU infrastructure, high performance computing, ML infrastructure, or large scale compute platforms. Experience with Kubernetes, cloud infrastructure, Linux, networking, and observability tooling. Excellent debugging, systems design, and operational problem solving skills. Strong communication skills and experience collaborating across engineering organizations. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
MLOps Engineer
Lumai Limited Oxford, Oxfordshire
The Opportunity Lumai is redefining how the world computes. We are an ambitious, venture-backed UK startup pioneering a breakthrough AI accelerator for data centers which uses 3D optical compute. Our radical technology uses light to perform computation at orders of magnitude faster speeds and at far greater scales than ever before, all whilst consuming far less energy than traditional approaches. Lumai is unlocking performance and efficiency gains that could transform the economics of AI and compute infrastructure and reshape how intelligence scales globally. If you are passionate about bringing groundbreaking technology to market, and want to be part of a team pushing the boundaries of what is physically possible, Lumai is where you can make it happen. About Lumai Founded in 2022, Lumai is a University of Oxford spinout using optical processing to accelerate large language models (LLMs) and other transformer-based AI systems. The team combines expertise in optical computing, machine learning, and physics. Lumai has already secured over $15 million in investment from leading deep-tech investors like Constructor Capital, IP Group, PhotonVentures and government grants, and is scaling rapidly to deploy the fastest optical compute currently available globally. The Role We are building custom AI hardware and the full-stack software ecosystem to run it. As our first dedicated MLOps Engineer, you will own the infrastructure that takes models from research to silicon-validated production - designing, building, and operating the pipelines, tooling, and platforms that let our AI and hardware teams move fast without breaking things. This is a high-impact, high-ownership role at the intersection of ML research, compiler stacks, and novel hardware. What You'll Do Design and operate end-to-end ML pipelines: data ingest, training, evaluation, quantisation, and deployment onto custom AI accelerator hardware Build and maintain experiment tracking, model registry, and versioning infrastructure (e.g. MLflow, W&B, or equivalent) tuned to our hardware-in-the-loop workflows Own CI/CD for ML: automated testing of model correctness, numerical accuracy, and on-chip performance after every change to models, compilers, or firmware Develop and maintain tooling for benchmarking model inference on custom silicon, including latency, throughput, power, and utilisation metrics Collaborate closely with ML researchers, compiler engineers, and hardware architects to identify and remove bottlenecks across the model-to-chip workflow Instrument and monitor production inference deployments; design alerting and rollback strategies appropriate to hardware-accelerated serving Manage compute resource scheduling across on-premises accelerator clusters and cloud (GPU/CPU) for training and simulation workloads Drive infrastructure-as-code practices: containerisation, orchestration (Kubernetes/Slurm), and reproducible environment management Contribute to the internal developer platform: self-service tooling, documentation, and runbooks that raise engineering productivity across the company What We're Looking For Must-Have 5+ years of software or infrastructure engineering experience, with at least 2 years in an ML or AI-adjacent role Strong Python skills and familiarity with major ML frameworks (PyTorch or JAX); comfortable reading and modifying model code Hands-on experience building and operating ML pipelines in production: data pipelines, training orchestration, evaluation, and serving Experience with experiment tracking and model lifecycle management tools (MLflow, W&B, DVC, or similar) Solid understanding of containerisation (Docker) and orchestration (Kubernetes or Slurm) for distributed compute workloads Infrastructure-as-code mindset: Terraform, Ansible, or equivalent; CI/CD pipelines (GitHub Actions, Jenkins, or similar) Experience with hardware-accelerated compute (CUDA/GPU workflows, profiling, performance tuning) - even if not on custom silicon Strong debugging and observability skills: distributed tracing, logging, metrics dashboards Ability to work effectively in a fast-moving, ambiguous environment where the hardware and software are both being built simultaneously Strong Preference For Experience with custom or novel accelerator hardware (FPGAs, ASICs, NPUs, or research chips) Familiarity with ML compiler stacks: MLIR, LLVM, TVM, XLA, or vendor-specific compilers (NVCC, TensorRT, etc.) Experience with model optimisation techniques: quantisation (INT8/INT4/FP8), pruning, distillation, or mixed-precision training Background in on-chip performance profiling and roofline analysis Exposure to chip bring-up workflows: running early software stacks on pre-silicon simulation or first-silicon hardware Contributions to open-source ML infrastructure or compiler tooling Experience in a deeptech, semiconductor, or hardware startup environment Compensation & Benefits Highly Competitive Salary: We are not saying our salary is a blank check, but let's just say it won't be a source of your stress Share Option Scheme: We are all in this together! We believe in shared success while we build the Lumai of tomorrow Pension Scheme: Plan for retirement with AVIVA Private Health Insurance: We firmly believe that you come first, and a happy you is a healthy you! Look after yourself and your loved ones with AXA Cycle to Work: Spread the cost of a bike, a bike and accessories or just accessories and save on tax L&D Allowance: Stay at the forefront of your field with a £500 annual development budget Subsidised On-site Lunches: Enjoy on-site healthy meals at half the price, as Lumai covers 50% of the cost Holidays: Enjoy some deserved "me time" with 25 days paid holiday (plus bank holidays) per year Socials: Be part of an inclusive community enjoying occasional all-company off-sites, lunches and socials Interview Process Our process is four stages. An initial conversation with our HR team to understand what you want from the role and what we want to Lumai is an equal opportunity employer. We make hiring decisions on merit, scope-fit, and the strength of the working relationship we expect to build with each hire. Applications welcome from candidates of any background. If you are not sure whether you are a fit, send a note anyway.
19/07/2026
Full time
The Opportunity Lumai is redefining how the world computes. We are an ambitious, venture-backed UK startup pioneering a breakthrough AI accelerator for data centers which uses 3D optical compute. Our radical technology uses light to perform computation at orders of magnitude faster speeds and at far greater scales than ever before, all whilst consuming far less energy than traditional approaches. Lumai is unlocking performance and efficiency gains that could transform the economics of AI and compute infrastructure and reshape how intelligence scales globally. If you are passionate about bringing groundbreaking technology to market, and want to be part of a team pushing the boundaries of what is physically possible, Lumai is where you can make it happen. About Lumai Founded in 2022, Lumai is a University of Oxford spinout using optical processing to accelerate large language models (LLMs) and other transformer-based AI systems. The team combines expertise in optical computing, machine learning, and physics. Lumai has already secured over $15 million in investment from leading deep-tech investors like Constructor Capital, IP Group, PhotonVentures and government grants, and is scaling rapidly to deploy the fastest optical compute currently available globally. The Role We are building custom AI hardware and the full-stack software ecosystem to run it. As our first dedicated MLOps Engineer, you will own the infrastructure that takes models from research to silicon-validated production - designing, building, and operating the pipelines, tooling, and platforms that let our AI and hardware teams move fast without breaking things. This is a high-impact, high-ownership role at the intersection of ML research, compiler stacks, and novel hardware. What You'll Do Design and operate end-to-end ML pipelines: data ingest, training, evaluation, quantisation, and deployment onto custom AI accelerator hardware Build and maintain experiment tracking, model registry, and versioning infrastructure (e.g. MLflow, W&B, or equivalent) tuned to our hardware-in-the-loop workflows Own CI/CD for ML: automated testing of model correctness, numerical accuracy, and on-chip performance after every change to models, compilers, or firmware Develop and maintain tooling for benchmarking model inference on custom silicon, including latency, throughput, power, and utilisation metrics Collaborate closely with ML researchers, compiler engineers, and hardware architects to identify and remove bottlenecks across the model-to-chip workflow Instrument and monitor production inference deployments; design alerting and rollback strategies appropriate to hardware-accelerated serving Manage compute resource scheduling across on-premises accelerator clusters and cloud (GPU/CPU) for training and simulation workloads Drive infrastructure-as-code practices: containerisation, orchestration (Kubernetes/Slurm), and reproducible environment management Contribute to the internal developer platform: self-service tooling, documentation, and runbooks that raise engineering productivity across the company What We're Looking For Must-Have 5+ years of software or infrastructure engineering experience, with at least 2 years in an ML or AI-adjacent role Strong Python skills and familiarity with major ML frameworks (PyTorch or JAX); comfortable reading and modifying model code Hands-on experience building and operating ML pipelines in production: data pipelines, training orchestration, evaluation, and serving Experience with experiment tracking and model lifecycle management tools (MLflow, W&B, DVC, or similar) Solid understanding of containerisation (Docker) and orchestration (Kubernetes or Slurm) for distributed compute workloads Infrastructure-as-code mindset: Terraform, Ansible, or equivalent; CI/CD pipelines (GitHub Actions, Jenkins, or similar) Experience with hardware-accelerated compute (CUDA/GPU workflows, profiling, performance tuning) - even if not on custom silicon Strong debugging and observability skills: distributed tracing, logging, metrics dashboards Ability to work effectively in a fast-moving, ambiguous environment where the hardware and software are both being built simultaneously Strong Preference For Experience with custom or novel accelerator hardware (FPGAs, ASICs, NPUs, or research chips) Familiarity with ML compiler stacks: MLIR, LLVM, TVM, XLA, or vendor-specific compilers (NVCC, TensorRT, etc.) Experience with model optimisation techniques: quantisation (INT8/INT4/FP8), pruning, distillation, or mixed-precision training Background in on-chip performance profiling and roofline analysis Exposure to chip bring-up workflows: running early software stacks on pre-silicon simulation or first-silicon hardware Contributions to open-source ML infrastructure or compiler tooling Experience in a deeptech, semiconductor, or hardware startup environment Compensation & Benefits Highly Competitive Salary: We are not saying our salary is a blank check, but let's just say it won't be a source of your stress Share Option Scheme: We are all in this together! We believe in shared success while we build the Lumai of tomorrow Pension Scheme: Plan for retirement with AVIVA Private Health Insurance: We firmly believe that you come first, and a happy you is a healthy you! Look after yourself and your loved ones with AXA Cycle to Work: Spread the cost of a bike, a bike and accessories or just accessories and save on tax L&D Allowance: Stay at the forefront of your field with a £500 annual development budget Subsidised On-site Lunches: Enjoy on-site healthy meals at half the price, as Lumai covers 50% of the cost Holidays: Enjoy some deserved "me time" with 25 days paid holiday (plus bank holidays) per year Socials: Be part of an inclusive community enjoying occasional all-company off-sites, lunches and socials Interview Process Our process is four stages. An initial conversation with our HR team to understand what you want from the role and what we want to Lumai is an equal opportunity employer. We make hiring decisions on merit, scope-fit, and the strength of the working relationship we expect to build with each hire. Applications welcome from candidates of any background. If you are not sure whether you are a fit, send a note anyway.
Software Engineer, General
AION
About aion Aion is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, Aion takes them from concept to production on a single, unified platform. We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters. We're a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we're building the next generation of enterprise AI and we're looking for exceptional people to help us scale. Who You Are You're a solid engineer with 2-4 years of experience building backend systems and platform infrastructure. You write clean, well-abstracted code with proper design patterns and comprehensive test coverage. You're comfortable working on both the Compute Platform (multi-cloud orchestration, resource management) and Inference Platform (model serving, autoscaling) under the guidance of senior engineers and platform leads. You have strong proficiency in Golang and understand how to build maintainable, production-grade distributed systems. You take pride in code quality, enjoy collaborating on low-level designs, and are eager to learn from experienced engineers while contributing meaningfully to critical infrastructure components. You're product-minded, you understand how your technical decisions impact developers using aion's platform and think about the end-to-end user experience. You're a team player comfortable wearing multiple hats one day you're building product features, the next you're joining customer calls to understand their deployment challenges, and the day after you're helping with UI/UX, customer success, documentation and product ops. What You'll Do Platform Development & Implementation Build and maintain platform services across aion's Compute and Inference platforms, working closely with senior engineers and platform leads Implement features for multi-cloud orchestration, resource scheduling, model deployment pipelines, and autoscaling systems Write well-maintained, production-grade code with proper abstractions, design patterns, and comprehensive test coverage Contribute to low-level design (LLD) including service APIs, database schema design, data models, and component interactions Collaborate with senior engineers on high-level design discussions, providing implementation perspectives and feasibility inputs Backend Systems & Distributed Infrastructure Develop RESTful APIs and gRPC services for platform control planes, resource management, and inference serving Design and implement database schemas for storing platform state, resource metadata, billing data, and observability metrics Work with distributed storage systems, message queues (Kafka, RabbitMQ), and databases (PostgreSQL, Redis) to build reliable platform components Build event-driven architectures for asynchronous processing, job scheduling, and platform automation Implement monitoring, logging, and alerting for platform services to ensure production reliability Code Quality & Engineering Excellence Write comprehensive unit tests, integration tests, and end-to-end tests to ensure code reliability Participate in code reviews, providing constructive feedback and learning from senior engineers' perspectives Refactor existing code to improve maintainability, performance, and scalability Document design decisions, API specifications, and operational runbooks for platform services Debug production issues and contribute to incident response and post-mortems Technical Skills & Experience 2-4 years of experience in backend engineering, platform development, or distributed systems Strong proficiency in Golang you write idiomatic Go code with proper error handling, concurrency patterns, and testing Solid understanding of backend systems fundamentals: RESTful APIs, microservices architecture, and API design principles Hands on experience with databases (PostgreSQL, MySQL) including schema design, query optimization, and transactions Familiarity with storage systems (object storage like S3, block storage, distributed file systems) and their use cases Experience working with message queues (Kafka, RabbitMQ, NATS) and event-driven architectures Understanding of distributed systems concepts: consensus, eventual consistency, fault tolerance, and retry mechanisms Experience with containerization (Docker) and basic Kubernetes concepts Knowledge of testing frameworks and practices (unit tests, integration tests, mocking) Familiarity with Git, CI/CD pipelines, and modern development workflows Exposure to cloud platforms (AWS/GCP/Azure) and their core services is a plus Experience with infrastructure-as-code (Terraform) or observability tools (Prometheus, Grafana) is beneficial Bonus/ Good to Have HPC & Cluster Management: Experience handling large-scale HPC clusters using Kubernetes and Slurm for job scheduling, resource allocation, and workload orchestration Data Engineering: Expertise with data pipelines, ETL systems, and large-scale data processing frameworks Systems-Level Programming: Experience with low-level systems programming such as storage systems, Kubernetes operators, OS-level software development, or daemon services (llm-d, system agents) ML Platform Engineering: Experience productionizing ML pipelines, batch job orchestration, model fine-tuning workflows, and Jupyter notebook orchestration systems Enterprise Deployment: Experience platformizing and packaging software for on-premises deployments or customer VPC installations with emphasis on security, compliance, and operational simplicity Preferred Attributes: High ownership, self driven and biased for action. Strong strategic thinking and ability to connect technical decisions to business impact. Excellent communication and mentoring skills. Thrives in ambiguity, fast-paced environments, and early stage startup culture. Why Join aion? Work directly with high pedigree founders shaping technical and product strategy. Build infrastructure powering the future of AI computers globally. Significant ownership and impact with equity reflective of your contributions. Competitive compensation, flexible work options, and wellness benefits
18/07/2026
Full time
About aion Aion is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, Aion takes them from concept to production on a single, unified platform. We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters. We're a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we're building the next generation of enterprise AI and we're looking for exceptional people to help us scale. Who You Are You're a solid engineer with 2-4 years of experience building backend systems and platform infrastructure. You write clean, well-abstracted code with proper design patterns and comprehensive test coverage. You're comfortable working on both the Compute Platform (multi-cloud orchestration, resource management) and Inference Platform (model serving, autoscaling) under the guidance of senior engineers and platform leads. You have strong proficiency in Golang and understand how to build maintainable, production-grade distributed systems. You take pride in code quality, enjoy collaborating on low-level designs, and are eager to learn from experienced engineers while contributing meaningfully to critical infrastructure components. You're product-minded, you understand how your technical decisions impact developers using aion's platform and think about the end-to-end user experience. You're a team player comfortable wearing multiple hats one day you're building product features, the next you're joining customer calls to understand their deployment challenges, and the day after you're helping with UI/UX, customer success, documentation and product ops. What You'll Do Platform Development & Implementation Build and maintain platform services across aion's Compute and Inference platforms, working closely with senior engineers and platform leads Implement features for multi-cloud orchestration, resource scheduling, model deployment pipelines, and autoscaling systems Write well-maintained, production-grade code with proper abstractions, design patterns, and comprehensive test coverage Contribute to low-level design (LLD) including service APIs, database schema design, data models, and component interactions Collaborate with senior engineers on high-level design discussions, providing implementation perspectives and feasibility inputs Backend Systems & Distributed Infrastructure Develop RESTful APIs and gRPC services for platform control planes, resource management, and inference serving Design and implement database schemas for storing platform state, resource metadata, billing data, and observability metrics Work with distributed storage systems, message queues (Kafka, RabbitMQ), and databases (PostgreSQL, Redis) to build reliable platform components Build event-driven architectures for asynchronous processing, job scheduling, and platform automation Implement monitoring, logging, and alerting for platform services to ensure production reliability Code Quality & Engineering Excellence Write comprehensive unit tests, integration tests, and end-to-end tests to ensure code reliability Participate in code reviews, providing constructive feedback and learning from senior engineers' perspectives Refactor existing code to improve maintainability, performance, and scalability Document design decisions, API specifications, and operational runbooks for platform services Debug production issues and contribute to incident response and post-mortems Technical Skills & Experience 2-4 years of experience in backend engineering, platform development, or distributed systems Strong proficiency in Golang you write idiomatic Go code with proper error handling, concurrency patterns, and testing Solid understanding of backend systems fundamentals: RESTful APIs, microservices architecture, and API design principles Hands on experience with databases (PostgreSQL, MySQL) including schema design, query optimization, and transactions Familiarity with storage systems (object storage like S3, block storage, distributed file systems) and their use cases Experience working with message queues (Kafka, RabbitMQ, NATS) and event-driven architectures Understanding of distributed systems concepts: consensus, eventual consistency, fault tolerance, and retry mechanisms Experience with containerization (Docker) and basic Kubernetes concepts Knowledge of testing frameworks and practices (unit tests, integration tests, mocking) Familiarity with Git, CI/CD pipelines, and modern development workflows Exposure to cloud platforms (AWS/GCP/Azure) and their core services is a plus Experience with infrastructure-as-code (Terraform) or observability tools (Prometheus, Grafana) is beneficial Bonus/ Good to Have HPC & Cluster Management: Experience handling large-scale HPC clusters using Kubernetes and Slurm for job scheduling, resource allocation, and workload orchestration Data Engineering: Expertise with data pipelines, ETL systems, and large-scale data processing frameworks Systems-Level Programming: Experience with low-level systems programming such as storage systems, Kubernetes operators, OS-level software development, or daemon services (llm-d, system agents) ML Platform Engineering: Experience productionizing ML pipelines, batch job orchestration, model fine-tuning workflows, and Jupyter notebook orchestration systems Enterprise Deployment: Experience platformizing and packaging software for on-premises deployments or customer VPC installations with emphasis on security, compliance, and operational simplicity Preferred Attributes: High ownership, self driven and biased for action. Strong strategic thinking and ability to connect technical decisions to business impact. Excellent communication and mentoring skills. Thrives in ambiguity, fast-paced environments, and early stage startup culture. Why Join aion? Work directly with high pedigree founders shaping technical and product strategy. Build infrastructure powering the future of AI computers globally. Significant ownership and impact with equity reflective of your contributions. Competitive compensation, flexible work options, and wellness benefits
Software Engineer, General
United States Digital Space LLC
About the companythe company is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, the company takes them from concept to production on a single, unified platform. We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters. We're a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we're building the next generation of enterprise AI and we're looking for exceptional people to help us scale. Who You Are You're a solid engineer with 2-4 years of experience building backend systems and platform infrastructure. You write clean, well-abstracted code with proper design patterns and comprehensive test coverage. You're comfortable working on both the Compute Platform (multi-cloud orchestration, resource management) and Inference Platform (model serving, autoscaling) under the guidance of senior engineers and platform leads. You have strong proficiency in Golang and understand how to build maintainable, production-grade distributed systems. You take pride in code quality, enjoy collaborating on low-level designs, and are eager to learn from experienced engineers while contributing meaningfully to critical infrastructure components. You're product minded, you understand how your technical decisions impact developers using the company's platform and think about the end to end user experience. You're a team player comfortable wearing multiple hats one day you're building product features, the next you're joining customer calls to understand their deployment challenges, and the day after you're helping with UI/UX, customer success, documentation and product ops. What You'll Do Platform Development & Implementation Build and maintain platform services across the company's Compute and Inference platforms, working closely with senior engineers and platform leads Implement features for multi cloud orchestration, resource scheduling, model deployment pipelines, and autoscaling systems Write well maintained, production grade code with proper abstractions, design patterns, and comprehensive test coverage Contribute to low level design (LLD) including service APIs, database schema design, data models, and component interactions Collaborate with senior engineers on high level design discussions, providing implementation perspectives and feasibility inputs Backend Systems & Distributed Infrastructure Develop RESTful APIs and gRPC services for platform control planes, resource management, and inference serving Design and implement database schemas for storing platform state, resource metadata, billing data, and observability metrics Work with distributed storage systems, message queues (Kafka, RabbitMQ), and databases (PostgreSQL, Redis) to build reliable platform components Build event driven architectures for asynchronous processing, job scheduling, and platform automation Implement monitoring, logging, and alerting for platform services to ensure production reliability Code Quality & Engineering Excellence Write comprehensive unit tests, integration tests, and end to end tests to ensure code reliability Participate in code reviews, providing constructive feedback and learning from senior engineers' perspectives Refactor existing code to improve maintainability, performance, and scalability Document design decisions, API specifications, and operational runbooks for platform services Debug production issues and contribute to incident response and post mortems. Requirements Technical Skills & Experience 2-4 years of experience in backend engineering, platform development, or distributed systems Strong proficiency in Golang you write idiomatic Go code with proper error handling, concurrency patterns, and testing Solid understanding of backend systems fundamentals: RESTful APIs, microservices architecture, and API design principles Hands on experience with databases (PostgreSQL, MySQL) including schema design, query optimization, and transactions Familiarity with storage systems (object storage like S3, block storage, distributed file systems) and their use cases Experience working with message queues (Kafka, RabbitMQ, NATS) and event driven architectures Understanding of distributed systems concepts: consensus, eventual consistency, fault tolerance, and retry mechanisms Experience with containerization (Docker) and basic Kubernetes concepts Knowledge of testing frameworks and practices (unit tests, integration tests, mocking) Familiarity with Git, CI/CD pipelines, and modern development workflows Exposure to cloud platforms (AWS/GCP/Azure) and their core services is a plus Experience with infrastructure as code (Terraform) or observability tools (Prometheus, Grafana) is beneficial Bonus/ Good to Have HPC & Cluster Management: Experience handling large-scale HPC clusters using Kubernetes and Slurm for job scheduling, resource allocation, and workload orchestration Data Engineering: Expertise with data pipelines, ETL systems, and large-scale data processing frameworks Systems Level Programming: Experience with low-level systems programming such as storage systems, Kubernetes operators, OS-level software development, or daemon services (llm d, system agents) ML Platform Engineering: Experience productionizing ML pipelines, batch job orchestration, model fine tuning workflows, and Jupyter notebook orchestration systems Enterprise Deployment: Experience platformizing and packaging software for on premises deployments or customer VPC installations with emphasis on security, compliance, and operational simplicity Benefits Preferred Attributes High ownership, self driven and biased for action. Strong strategic thinking and ability to connect technical decisions to business impact. Excellent communication and mentoring skills. Thrives in ambiguity, fast paced environments, and early stage startup culture. Why Join the company? Work directly with high pedigree founders shaping technical and product strategy. Build infrastructure powering the future of AI computers globally. Significant ownership and impact with equity reflective of your contributions. Competitive compensation, flexible work options, and wellness benefits.
16/07/2026
Full time
About the companythe company is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, the company takes them from concept to production on a single, unified platform. We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters. We're a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we're building the next generation of enterprise AI and we're looking for exceptional people to help us scale. Who You Are You're a solid engineer with 2-4 years of experience building backend systems and platform infrastructure. You write clean, well-abstracted code with proper design patterns and comprehensive test coverage. You're comfortable working on both the Compute Platform (multi-cloud orchestration, resource management) and Inference Platform (model serving, autoscaling) under the guidance of senior engineers and platform leads. You have strong proficiency in Golang and understand how to build maintainable, production-grade distributed systems. You take pride in code quality, enjoy collaborating on low-level designs, and are eager to learn from experienced engineers while contributing meaningfully to critical infrastructure components. You're product minded, you understand how your technical decisions impact developers using the company's platform and think about the end to end user experience. You're a team player comfortable wearing multiple hats one day you're building product features, the next you're joining customer calls to understand their deployment challenges, and the day after you're helping with UI/UX, customer success, documentation and product ops. What You'll Do Platform Development & Implementation Build and maintain platform services across the company's Compute and Inference platforms, working closely with senior engineers and platform leads Implement features for multi cloud orchestration, resource scheduling, model deployment pipelines, and autoscaling systems Write well maintained, production grade code with proper abstractions, design patterns, and comprehensive test coverage Contribute to low level design (LLD) including service APIs, database schema design, data models, and component interactions Collaborate with senior engineers on high level design discussions, providing implementation perspectives and feasibility inputs Backend Systems & Distributed Infrastructure Develop RESTful APIs and gRPC services for platform control planes, resource management, and inference serving Design and implement database schemas for storing platform state, resource metadata, billing data, and observability metrics Work with distributed storage systems, message queues (Kafka, RabbitMQ), and databases (PostgreSQL, Redis) to build reliable platform components Build event driven architectures for asynchronous processing, job scheduling, and platform automation Implement monitoring, logging, and alerting for platform services to ensure production reliability Code Quality & Engineering Excellence Write comprehensive unit tests, integration tests, and end to end tests to ensure code reliability Participate in code reviews, providing constructive feedback and learning from senior engineers' perspectives Refactor existing code to improve maintainability, performance, and scalability Document design decisions, API specifications, and operational runbooks for platform services Debug production issues and contribute to incident response and post mortems. Requirements Technical Skills & Experience 2-4 years of experience in backend engineering, platform development, or distributed systems Strong proficiency in Golang you write idiomatic Go code with proper error handling, concurrency patterns, and testing Solid understanding of backend systems fundamentals: RESTful APIs, microservices architecture, and API design principles Hands on experience with databases (PostgreSQL, MySQL) including schema design, query optimization, and transactions Familiarity with storage systems (object storage like S3, block storage, distributed file systems) and their use cases Experience working with message queues (Kafka, RabbitMQ, NATS) and event driven architectures Understanding of distributed systems concepts: consensus, eventual consistency, fault tolerance, and retry mechanisms Experience with containerization (Docker) and basic Kubernetes concepts Knowledge of testing frameworks and practices (unit tests, integration tests, mocking) Familiarity with Git, CI/CD pipelines, and modern development workflows Exposure to cloud platforms (AWS/GCP/Azure) and their core services is a plus Experience with infrastructure as code (Terraform) or observability tools (Prometheus, Grafana) is beneficial Bonus/ Good to Have HPC & Cluster Management: Experience handling large-scale HPC clusters using Kubernetes and Slurm for job scheduling, resource allocation, and workload orchestration Data Engineering: Expertise with data pipelines, ETL systems, and large-scale data processing frameworks Systems Level Programming: Experience with low-level systems programming such as storage systems, Kubernetes operators, OS-level software development, or daemon services (llm d, system agents) ML Platform Engineering: Experience productionizing ML pipelines, batch job orchestration, model fine tuning workflows, and Jupyter notebook orchestration systems Enterprise Deployment: Experience platformizing and packaging software for on premises deployments or customer VPC installations with emphasis on security, compliance, and operational simplicity Benefits Preferred Attributes High ownership, self driven and biased for action. Strong strategic thinking and ability to connect technical decisions to business impact. Excellent communication and mentoring skills. Thrives in ambiguity, fast paced environments, and early stage startup culture. Why Join the company? Work directly with high pedigree founders shaping technical and product strategy. Build infrastructure powering the future of AI computers globally. Significant ownership and impact with equity reflective of your contributions. Competitive compensation, flexible work options, and wellness benefits.
Senior Software Engineer, Inference Platform
United States Digital Space LLC
About the companythe company is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, the company takes them from concept to production on a single, unified platform. We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters. We're a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we're building the next generation of enterprise AI and we're looking for exceptional people to help us scale. Who You Are You're a seasoned engineer who has built and scaled high-performance inference systems for AI/ML workloads. You understand the complexities of serving models at scale latency optimization, resource orchestration, autoscaling dynamics, and production reliability. You've designed distributed systems that handle thousands of requests per second while maintaining sub second response times and cost efficiency. Experience with Golang is strongly preferred, and exposure to inference engines (vLLM, TGI, TensorRT), containerization, and distributed systems is an added bonus. You take ownership of platform level decisions, think strategically about performance vs. cost trade offs, and want your work to power AI inference for thousands of developers globally. You're product minded, you understand how your technical decisions impact developers using the company's platform and think about the end to end user experience. You're a team player comfortable wearing multiple hats one day you're optimizing inference latency, the next you're joining customer calls to understand their deployment challenges, and the day after you're helping with UI/UX, customer success, documentation and product ops. What You'll Do Inference Platform Architecture & Core Services Design and build the company's inference service platform the backbone for serving AI models at scale across diverse workloads Own and architect core platform components: AI Gateway, Resource Orchestrator, Runtime Engines, and Autoscaler Design highly modular, scalable, and extensible low level designs (LLDs) for inference infrastructure components Lead high level design discussions, establish architectural patterns, and drive technical decision making for the inference stack Model Deployment & Lifecycle Management Understand and optimize the dynamics of model deployment, version upgrades, and rollback strategies Build robust deployment pipelines for seamless model updates with zero downtime deployments Design intelligent routing systems for multi model serving, A/B testing, and canary deployments Implement strategies for efficient GPU utilization and model cold start optimization Performance & Distributed Systems Implement highly performant and optimized software for low latency, high throughput inference serving Build and debug production grade code in distributed systems handling real time AI workloads Optimize inference pipelines for latency, throughput, batching efficiency, and resource utilization Design fault tolerant systems with graceful degradation and automatic recovery mechanisms Observability & Engineering Excellence Build high performance telemetry and observability stack for inference metrics, performance tracking, and debugging Implement comprehensive monitoring for model latency, throughput, error rates, GPU utilization, and cost per inference Conduct thorough code reviews to maintain code quality, performance standards, and architectural consistency Establish engineering best practices for testing, documentation, and production readiness. Requirements Technical Skills & Experience 4+ years of experience building and scaling backend systems, distributed platforms, or inference infrastructure Strong understanding of AI/ML inference systems and experience with inference engines (vLLM, TGI, TensorRT LLM, or similar) Deep knowledge of distributed systems design, microservices architecture, and API gateway patterns Proficiency in Golang strongly preferred; Python, Rust, C++ for performance critical components a plus Experience with container orchestration (Kubernetes, Docker) and infrastructure as code Solid understanding of autoscaling strategies, load balancing, and resource scheduling algorithms Experience building high throughput, low latency systems with sub 100ms response time requirements Familiarity with message queues (Kafka, RabbitMQ), databases (PostgreSQL, Redis), and event driven architectures Knowledge of GPU computing, model serving optimizations (batching, quantization, multi tenancy), and resource allocation Experience with observability tools (Prometheus, Grafana, OpenTelemetry) and distributed tracing Understanding of API design, rate limiting, authentication/authorization, and security best practices Exposure to AI model deployment workflows and model lifecycle management is highly desirable Bonus / Good to Have HPC & Cluster Management: Experience handling large scale HPC clusters using Kubernetes and Slurm for job scheduling, resource allocation, and workload orchestration Data Engineering: Expertise with data pipelines, ETL systems, and large scale data processing frameworks Systems-Level Programming: Experience with low level systems programming such as storage systems, Kubernetes operators, OS-level software development, or daemon services (llm d, system agents) ML Platform Engineering: Experience productionizing ML pipelines, batch job orchestration, model fine tuning workflows, and Jupyter notebook orchestration systems Enterprise Deployment: Experience platformizing and packaging software for on premises deployments or customer VPC installations with emphasis on security, compliance, and operational simplicity Benefits Preferred Attributes High ownership, self driven and biased for action. Strong strategic thinking and ability to connect technical decisions to business impact. Excellent communication and mentoring skills. Thrives in ambiguity, fast paced environments, and early-stage startup culture. Why Join the company? Work directly with high pedigree founders shaping technical and product strategy. Build infrastructure powering the future of AI computers globally. Significant ownership and impact with equity reflective of your contributions. Competitive compensation, flexible work options, and wellness benefits
16/07/2026
Full time
About the companythe company is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, the company takes them from concept to production on a single, unified platform. We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters. We're a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we're building the next generation of enterprise AI and we're looking for exceptional people to help us scale. Who You Are You're a seasoned engineer who has built and scaled high-performance inference systems for AI/ML workloads. You understand the complexities of serving models at scale latency optimization, resource orchestration, autoscaling dynamics, and production reliability. You've designed distributed systems that handle thousands of requests per second while maintaining sub second response times and cost efficiency. Experience with Golang is strongly preferred, and exposure to inference engines (vLLM, TGI, TensorRT), containerization, and distributed systems is an added bonus. You take ownership of platform level decisions, think strategically about performance vs. cost trade offs, and want your work to power AI inference for thousands of developers globally. You're product minded, you understand how your technical decisions impact developers using the company's platform and think about the end to end user experience. You're a team player comfortable wearing multiple hats one day you're optimizing inference latency, the next you're joining customer calls to understand their deployment challenges, and the day after you're helping with UI/UX, customer success, documentation and product ops. What You'll Do Inference Platform Architecture & Core Services Design and build the company's inference service platform the backbone for serving AI models at scale across diverse workloads Own and architect core platform components: AI Gateway, Resource Orchestrator, Runtime Engines, and Autoscaler Design highly modular, scalable, and extensible low level designs (LLDs) for inference infrastructure components Lead high level design discussions, establish architectural patterns, and drive technical decision making for the inference stack Model Deployment & Lifecycle Management Understand and optimize the dynamics of model deployment, version upgrades, and rollback strategies Build robust deployment pipelines for seamless model updates with zero downtime deployments Design intelligent routing systems for multi model serving, A/B testing, and canary deployments Implement strategies for efficient GPU utilization and model cold start optimization Performance & Distributed Systems Implement highly performant and optimized software for low latency, high throughput inference serving Build and debug production grade code in distributed systems handling real time AI workloads Optimize inference pipelines for latency, throughput, batching efficiency, and resource utilization Design fault tolerant systems with graceful degradation and automatic recovery mechanisms Observability & Engineering Excellence Build high performance telemetry and observability stack for inference metrics, performance tracking, and debugging Implement comprehensive monitoring for model latency, throughput, error rates, GPU utilization, and cost per inference Conduct thorough code reviews to maintain code quality, performance standards, and architectural consistency Establish engineering best practices for testing, documentation, and production readiness. Requirements Technical Skills & Experience 4+ years of experience building and scaling backend systems, distributed platforms, or inference infrastructure Strong understanding of AI/ML inference systems and experience with inference engines (vLLM, TGI, TensorRT LLM, or similar) Deep knowledge of distributed systems design, microservices architecture, and API gateway patterns Proficiency in Golang strongly preferred; Python, Rust, C++ for performance critical components a plus Experience with container orchestration (Kubernetes, Docker) and infrastructure as code Solid understanding of autoscaling strategies, load balancing, and resource scheduling algorithms Experience building high throughput, low latency systems with sub 100ms response time requirements Familiarity with message queues (Kafka, RabbitMQ), databases (PostgreSQL, Redis), and event driven architectures Knowledge of GPU computing, model serving optimizations (batching, quantization, multi tenancy), and resource allocation Experience with observability tools (Prometheus, Grafana, OpenTelemetry) and distributed tracing Understanding of API design, rate limiting, authentication/authorization, and security best practices Exposure to AI model deployment workflows and model lifecycle management is highly desirable Bonus / Good to Have HPC & Cluster Management: Experience handling large scale HPC clusters using Kubernetes and Slurm for job scheduling, resource allocation, and workload orchestration Data Engineering: Expertise with data pipelines, ETL systems, and large scale data processing frameworks Systems-Level Programming: Experience with low level systems programming such as storage systems, Kubernetes operators, OS-level software development, or daemon services (llm d, system agents) ML Platform Engineering: Experience productionizing ML pipelines, batch job orchestration, model fine tuning workflows, and Jupyter notebook orchestration systems Enterprise Deployment: Experience platformizing and packaging software for on premises deployments or customer VPC installations with emphasis on security, compliance, and operational simplicity Benefits Preferred Attributes High ownership, self driven and biased for action. Strong strategic thinking and ability to connect technical decisions to business impact. Excellent communication and mentoring skills. Thrives in ambiguity, fast paced environments, and early-stage startup culture. Why Join the company? Work directly with high pedigree founders shaping technical and product strategy. Build infrastructure powering the future of AI computers globally. Significant ownership and impact with equity reflective of your contributions. Competitive compensation, flexible work options, and wellness benefits
Software Engineer, GPU Infrastructure- ChatGPT Engineering
The Consulting Solutions
About the Team ChatGPT Engineering builds and operates the compute platform powering one of the world's largest AI products. Every ChatGPT conversation relies on massive GPU clusters serving inference workloads with high reliability, efficiency, and performance. As our GPU fleet continues to grow, we're investing in the infrastructure that operates it. Our team builds the tooling, automation, and intelligent systems that make GPU infrastructure scalable, observable, and increasingly autonomous. We work across production engineering, distributed systems, capacity management, and AI powered operational tooling to help researchers and product teams move faster while maximizing efficiency of every GPU. This is a unique opportunity to work on infrastructure at the frontier of AI, where small improvements in fleet efficiency, reliability, and automation have an outsized impact on the development and deployment of AGI. About the Role We're looking for a Software Engineer with deep experience operating large scale GPU or compute infrastructure. You'll design and build the systems that manage GPU clusters at scale-from fleet health and capacity planning to operational automation and intelligent agents that reduce manual intervention. You'll partner closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and overall compute utilization. This role is ideal for engineers who enjoy solving complex operational challenges, building internal platforms, and working on infrastructure that directly powers frontier AI. In This Role, You Will Design, build, and operate software that manages large scale GPU infrastructure supporting ChatGPT inference. Build internal platforms, tooling, and AI powered agents that automate fleet operations and reduce operational overhead. Improve observability, reliability, and operational efficiency across thousands of GPUs. Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response. Identify infrastructure bottlenecks and implement solutions that improve utilization, scalability, and performance. Partner closely with research, platform, networking, and systems teams to continuously improve our compute platform. Help establish engineering best practices around operational excellence, automation, and infrastructure reliability. You Might Thrive in This Role If You Have experience operating large scale production infrastructure, preferably GPU clusters or other compute intensive distributed systems. Have a background in Production Engineering, Site Reliability Engineering (SRE), Infrastructure Engineering, or Platform Engineering. Have built software that automates operational workflows rather than relying on manual processes. Have experience with Kubernetes, Linux systems, container orchestration, or distributed infrastructure. Understand infrastructure observability, monitoring, capacity planning, and incident management. Enjoy identifying cross team pain points and building reusable platforms that improve developer productivity. Are comfortable working across software engineering and systems operations, owning problems end to end. Thrive in fast moving environments with significant technical ambiguity. Qualifications 5+ years of software engineering experience building production infrastructure. Strong programming skills in Go, Python, C++, Rust, or similar systems languages. Experience designing and operating highly available distributed systems. Experience with GPU infrastructure, high performance computing, ML infrastructure, or large scale compute platforms. Experience with Kubernetes, cloud infrastructure, Linux, networking, and observability tooling. Excellent debugging, systems design, and operational problem solving skills. Strong communication skills and experience collaborating across engineering organizations. EEO Statement We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.
12/07/2026
Full time
About the Team ChatGPT Engineering builds and operates the compute platform powering one of the world's largest AI products. Every ChatGPT conversation relies on massive GPU clusters serving inference workloads with high reliability, efficiency, and performance. As our GPU fleet continues to grow, we're investing in the infrastructure that operates it. Our team builds the tooling, automation, and intelligent systems that make GPU infrastructure scalable, observable, and increasingly autonomous. We work across production engineering, distributed systems, capacity management, and AI powered operational tooling to help researchers and product teams move faster while maximizing efficiency of every GPU. This is a unique opportunity to work on infrastructure at the frontier of AI, where small improvements in fleet efficiency, reliability, and automation have an outsized impact on the development and deployment of AGI. About the Role We're looking for a Software Engineer with deep experience operating large scale GPU or compute infrastructure. You'll design and build the systems that manage GPU clusters at scale-from fleet health and capacity planning to operational automation and intelligent agents that reduce manual intervention. You'll partner closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and overall compute utilization. This role is ideal for engineers who enjoy solving complex operational challenges, building internal platforms, and working on infrastructure that directly powers frontier AI. In This Role, You Will Design, build, and operate software that manages large scale GPU infrastructure supporting ChatGPT inference. Build internal platforms, tooling, and AI powered agents that automate fleet operations and reduce operational overhead. Improve observability, reliability, and operational efficiency across thousands of GPUs. Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response. Identify infrastructure bottlenecks and implement solutions that improve utilization, scalability, and performance. Partner closely with research, platform, networking, and systems teams to continuously improve our compute platform. Help establish engineering best practices around operational excellence, automation, and infrastructure reliability. You Might Thrive in This Role If You Have experience operating large scale production infrastructure, preferably GPU clusters or other compute intensive distributed systems. Have a background in Production Engineering, Site Reliability Engineering (SRE), Infrastructure Engineering, or Platform Engineering. Have built software that automates operational workflows rather than relying on manual processes. Have experience with Kubernetes, Linux systems, container orchestration, or distributed infrastructure. Understand infrastructure observability, monitoring, capacity planning, and incident management. Enjoy identifying cross team pain points and building reusable platforms that improve developer productivity. Are comfortable working across software engineering and systems operations, owning problems end to end. Thrive in fast moving environments with significant technical ambiguity. Qualifications 5+ years of software engineering experience building production infrastructure. Strong programming skills in Go, Python, C++, Rust, or similar systems languages. Experience designing and operating highly available distributed systems. Experience with GPU infrastructure, high performance computing, ML infrastructure, or large scale compute platforms. Experience with Kubernetes, cloud infrastructure, Linux, networking, and observability tooling. Excellent debugging, systems design, and operational problem solving skills. Strong communication skills and experience collaborating across engineering organizations. EEO Statement We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.
Software Engineer, GPU Infrastructure- ChatGPT Engineering
OpenAI
About the Team ChatGPT Engineering builds and operates the compute platform powering one of the world's largest AI products. Every ChatGPT conversation relies on massive GPU clusters serving inference workloads with high reliability, efficiency, and performance. As our GPU fleet continues to grow, we're investing in the infrastructure that operates it. Our team builds the tooling, automation, and intelligent systems that make GPU infrastructure scalable, observable, and increasingly autonomous. We work across production engineering, distributed systems, capacity management, and AI-powered operational tooling to help researchers and product teams move faster while maximizing the efficiency of every GPU. This is a unique opportunity to work on infrastructure at the frontier of AI, where small improvements in fleet efficiency, reliability, and automation have an outsized impact on the development and deployment of AGI. About the Role We're looking for a Software Engineer with deep experience operating large-scale GPU or compute infrastructure. You'll design and build the systems that manage GPU clusters at scale-from fleet health and capacity planning to operational automation and intelligent agents that reduce manual intervention. You'll partner closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and overall compute utilization. This role is ideal for engineers who enjoy solving complex operational challenges, building internal platforms, and working on infrastructure that directly powers frontier AI. In This Role, You Will Design, build, and operate software that manages large-scale GPU infrastructure supporting ChatGPT inference. Build internal platforms, tooling, and AI-powered agents that automate fleet operations and reduce operational overhead. Improve observability, reliability, and operational efficiency across thousands of GPUs. Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response. Identify infrastructure bottlenecks and implement solutions that improve utilization, scalability, and performance. Partner closely with research, platform, networking, and systems teams to continuously improve our compute platform. Help establish engineering best practices around operational excellence, automation, and infrastructure reliability. You Might Thrive in This Role If You Have experience operating large-scale production infrastructure, preferably GPU clusters or other compute-intensive distributed systems. Have a background in Production Engineering, Site Reliability Engineering (SRE), Infrastructure Engineering, or Platform Engineering. Have built software that automates operational workflows rather than relying on manual processes. Have experience with Kubernetes, Linux systems, container orchestration, or distributed infrastructure. Understand infrastructure observability, monitoring, capacity planning, and incident management. Enjoy identifying cross-team pain points and building reusable platforms that improve developer productivity. Are comfortable working across software engineering and systems operations, owning problems end-to-end. Thrive in fast-moving environments with significant technical ambiguity. Qualifications 5+ years of software engineering experience building production infrastructure. Strong programming skills in Go, Python, C++, Rust, or similar systems languages. Experience designing and operating highly available distributed systems. Experience with GPU infrastructure, high-performance computing, ML infrastructure, or large-scale compute platforms. Experience with Kubernetes, cloud infrastructure, Linux, networking, and observability tooling. Excellent debugging, systems design, and operational problem-solving skills. Strong communication skills and experience collaborating across engineering organizations. Equal Opportunity Employment We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. Legal Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations. To notify OpenAI that you believe this job posting is non compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance. We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link. OpenAI Global Applicant Privacy Policy At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.
12/07/2026
Full time
About the Team ChatGPT Engineering builds and operates the compute platform powering one of the world's largest AI products. Every ChatGPT conversation relies on massive GPU clusters serving inference workloads with high reliability, efficiency, and performance. As our GPU fleet continues to grow, we're investing in the infrastructure that operates it. Our team builds the tooling, automation, and intelligent systems that make GPU infrastructure scalable, observable, and increasingly autonomous. We work across production engineering, distributed systems, capacity management, and AI-powered operational tooling to help researchers and product teams move faster while maximizing the efficiency of every GPU. This is a unique opportunity to work on infrastructure at the frontier of AI, where small improvements in fleet efficiency, reliability, and automation have an outsized impact on the development and deployment of AGI. About the Role We're looking for a Software Engineer with deep experience operating large-scale GPU or compute infrastructure. You'll design and build the systems that manage GPU clusters at scale-from fleet health and capacity planning to operational automation and intelligent agents that reduce manual intervention. You'll partner closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and overall compute utilization. This role is ideal for engineers who enjoy solving complex operational challenges, building internal platforms, and working on infrastructure that directly powers frontier AI. In This Role, You Will Design, build, and operate software that manages large-scale GPU infrastructure supporting ChatGPT inference. Build internal platforms, tooling, and AI-powered agents that automate fleet operations and reduce operational overhead. Improve observability, reliability, and operational efficiency across thousands of GPUs. Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response. Identify infrastructure bottlenecks and implement solutions that improve utilization, scalability, and performance. Partner closely with research, platform, networking, and systems teams to continuously improve our compute platform. Help establish engineering best practices around operational excellence, automation, and infrastructure reliability. You Might Thrive in This Role If You Have experience operating large-scale production infrastructure, preferably GPU clusters or other compute-intensive distributed systems. Have a background in Production Engineering, Site Reliability Engineering (SRE), Infrastructure Engineering, or Platform Engineering. Have built software that automates operational workflows rather than relying on manual processes. Have experience with Kubernetes, Linux systems, container orchestration, or distributed infrastructure. Understand infrastructure observability, monitoring, capacity planning, and incident management. Enjoy identifying cross-team pain points and building reusable platforms that improve developer productivity. Are comfortable working across software engineering and systems operations, owning problems end-to-end. Thrive in fast-moving environments with significant technical ambiguity. Qualifications 5+ years of software engineering experience building production infrastructure. Strong programming skills in Go, Python, C++, Rust, or similar systems languages. Experience designing and operating highly available distributed systems. Experience with GPU infrastructure, high-performance computing, ML infrastructure, or large-scale compute platforms. Experience with Kubernetes, cloud infrastructure, Linux, networking, and observability tooling. Excellent debugging, systems design, and operational problem-solving skills. Strong communication skills and experience collaborating across engineering organizations. Equal Opportunity Employment We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. Legal Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations. To notify OpenAI that you believe this job posting is non compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance. We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link. OpenAI Global Applicant Privacy Policy At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.
Software Developer - Distributed Scheduling & Workload Orchestration
Viridiengroup Crawley, Sussex
) for more information.Viridien () is an advanced technology, digital and Earth data company that pushes the boundaries of science for a more prosperous and sustainable future. With our ingenuity, drive and deep curiosity we discover new insights, innovations, and solutions that efficiently and responsibly resolve complex natural resource, digital, energy transition and infrastructure challenges. Job Details Viridien is seeking a Software Developer - Distributed Scheduling & Workload Orchestration to design, build, and improve systems responsible for job scheduling, resource allocation, and workload orchestration across distributed environments.This role focuses on building scalable and reliable systems that coordinate workloads across clusters, using technologies such as Slurm, Golang, Java, PostgreSQL, and containerised microservices. About The Team You will join a team working on distributed systems and infrastructure that support large-scale compute and workload execution.The team focuses on building reliable scheduling and orchestration systems that manage resources efficiently across complex environments. Key Responsibilities - Scheduling & Orchestration Design and develop systems for job scheduling and workload orchestration. Integrate and extend scheduling capabilities using tools such as Slurm. Manage job lifecycles, resource allocation, and execution workflows. - Backend & Data Systems Design and build APIs and backend services supporting scheduling systems. Work with PostgreSQL to manage system state and coordination. - Performance & Reliability Analyse and improve system performance, scalability, and reliability. Ensure efficient resource utilisation across distributed environments. - Architecture & Collaboration Participate in system design and architecture discussions. Work with cross-functional teams to evolve scheduling and orchestration capabilities. Qualifications Required Strong software development experience. Proven experience building backend services or distributed systems. Experience with job scheduling, orchestration systems, or resource management concepts. Strong understanding of distributed systems concepts such as coordination, consistency, and fault tolerance. Experience working with PostgreSQL. Experience designing APIs and backend services. Familiarity with containerised environments and microservices architectures. Strong problem-solving and analytical skills. Preferred Experience with Slurm or similar workload managers. Experience in HPC or large-scale compute environments. Experience with Golang or Java. Familiarity with C/C++ and performance-critical systems. Experience providing technical or project leadership. Competitive salary commensurate with experience Highly attractive bonus scheme Initial 22 days annual leave with future increases, complemented by a flexible buying and selling holiday program Company pension with generous employer contribution Wellbeing Unmind app - puts you in control of your mental health A flexible benefits platform with numerous discount schemes - gym membership, restaurants, cinema tickets, and much more! Regular social club events, spontaneous reward events throughout the year Cycle purchase scheme Flexible Private Medical & Dental care programmes Sponsorship of visas/comprehensive relocation packages Bank Holiday Swap - our holiday swap program allows you to change it for another day of your choice! Relaxed dress code policy Learning and Development At Viridien, we foster a culture of continuous learning and provide tailored training programs through our Learning Hub, designed to enhance technical, commercial, and personal growth. We Care About The Environment We encourage and actively support a strong sense of community, through volunteering and various company initiatives, as well as a strong company commitment to protecting our environment through sustainable solutions, energy saving and waste reduction enterprises. Our Hiring Process At Viridien, we are committed to delivering a respectful, inclusive, and transparent recruitment experience.Due to the high volume of applications we receive, we may not be able to provide individual feedback to every applicant. Only candidates whose qualifications closely match the role criteria will be contacted for an interview. We do, however, aim to share personalized feedback with those who progress to the first round of interviews and beyond.We are also dedicated to ensuring that our hiring process accessible to all. If you require any reasonable adjustments to fully participate in the application or interview stages, please don't hesitate to contact your recruiter directly.We see things differently. Diversity fuels our innovation, we value the unique ways in which we differ, and we are committed to equal employment opportunities for all professionals.Create a brighter future for
12/07/2026
Full time
) for more information.Viridien () is an advanced technology, digital and Earth data company that pushes the boundaries of science for a more prosperous and sustainable future. With our ingenuity, drive and deep curiosity we discover new insights, innovations, and solutions that efficiently and responsibly resolve complex natural resource, digital, energy transition and infrastructure challenges. Job Details Viridien is seeking a Software Developer - Distributed Scheduling & Workload Orchestration to design, build, and improve systems responsible for job scheduling, resource allocation, and workload orchestration across distributed environments.This role focuses on building scalable and reliable systems that coordinate workloads across clusters, using technologies such as Slurm, Golang, Java, PostgreSQL, and containerised microservices. About The Team You will join a team working on distributed systems and infrastructure that support large-scale compute and workload execution.The team focuses on building reliable scheduling and orchestration systems that manage resources efficiently across complex environments. Key Responsibilities - Scheduling & Orchestration Design and develop systems for job scheduling and workload orchestration. Integrate and extend scheduling capabilities using tools such as Slurm. Manage job lifecycles, resource allocation, and execution workflows. - Backend & Data Systems Design and build APIs and backend services supporting scheduling systems. Work with PostgreSQL to manage system state and coordination. - Performance & Reliability Analyse and improve system performance, scalability, and reliability. Ensure efficient resource utilisation across distributed environments. - Architecture & Collaboration Participate in system design and architecture discussions. Work with cross-functional teams to evolve scheduling and orchestration capabilities. Qualifications Required Strong software development experience. Proven experience building backend services or distributed systems. Experience with job scheduling, orchestration systems, or resource management concepts. Strong understanding of distributed systems concepts such as coordination, consistency, and fault tolerance. Experience working with PostgreSQL. Experience designing APIs and backend services. Familiarity with containerised environments and microservices architectures. Strong problem-solving and analytical skills. Preferred Experience with Slurm or similar workload managers. Experience in HPC or large-scale compute environments. Experience with Golang or Java. Familiarity with C/C++ and performance-critical systems. Experience providing technical or project leadership. Competitive salary commensurate with experience Highly attractive bonus scheme Initial 22 days annual leave with future increases, complemented by a flexible buying and selling holiday program Company pension with generous employer contribution Wellbeing Unmind app - puts you in control of your mental health A flexible benefits platform with numerous discount schemes - gym membership, restaurants, cinema tickets, and much more! Regular social club events, spontaneous reward events throughout the year Cycle purchase scheme Flexible Private Medical & Dental care programmes Sponsorship of visas/comprehensive relocation packages Bank Holiday Swap - our holiday swap program allows you to change it for another day of your choice! Relaxed dress code policy Learning and Development At Viridien, we foster a culture of continuous learning and provide tailored training programs through our Learning Hub, designed to enhance technical, commercial, and personal growth. We Care About The Environment We encourage and actively support a strong sense of community, through volunteering and various company initiatives, as well as a strong company commitment to protecting our environment through sustainable solutions, energy saving and waste reduction enterprises. Our Hiring Process At Viridien, we are committed to delivering a respectful, inclusive, and transparent recruitment experience.Due to the high volume of applications we receive, we may not be able to provide individual feedback to every applicant. Only candidates whose qualifications closely match the role criteria will be contacted for an interview. We do, however, aim to share personalized feedback with those who progress to the first round of interviews and beyond.We are also dedicated to ensuring that our hiring process accessible to all. If you require any reasonable adjustments to fully participate in the application or interview stages, please don't hesitate to contact your recruiter directly.We see things differently. Diversity fuels our innovation, we value the unique ways in which we differ, and we are committed to equal employment opportunities for all professionals.Create a brighter future for
Distributed Scheduling & Workload Orchestration Engineer
Viridiengroup Crawley, Sussex
A technology and data company in Crawley seeks a Software Developer focusing on distributed scheduling and workload orchestration. This position requires strong software development skills and experience with backend services and job scheduling systems. The developer will design systems for workload orchestration using tools such as Slurm, PostgreSQL, and microservices. The company offers competitive salary, bonuses, and flexible benefits, creating a supportive and innovative work environment.
12/07/2026
Full time
A technology and data company in Crawley seeks a Software Developer focusing on distributed scheduling and workload orchestration. This position requires strong software development skills and experience with backend services and job scheduling systems. The developer will design systems for workload orchestration using tools such as Slurm, PostgreSQL, and microservices. The company offers competitive salary, bonuses, and flexible benefits, creating a supportive and innovative work environment.
Software Developer - Distributed Scheduling & Workload Orchestration
CGG Services (UK) Limited Crawley, Sussex
Software Developer - Distributed Scheduling & Workload Orchestration This role focuses on building scalable and reliable systems that coordinate workloads across clusters, using technologies such as Slurm, Golang, Java, PostgreSQL, and containerised microservices. Responsibilities Design and develop systems for job scheduling and workload orchestration; integrate and extend scheduling capabilities using tools such as Slurm; manage job lifecycles, resource allocation, and execution workflows. Design and build APIs and backend services supporting scheduling systems; work with PostgreSQL to manage system state and coordination. Analyse and improve system performance, scalability, and reliability; ensure efficient resource utilisation across distributed environments. Participate in system design and architecture discussions; collaborate with cross functional teams to evolve scheduling and orchestration capabilities. Qualifications Strong software development experience. Proven experience building backend services or distributed systems. Experience with job scheduling, orchestration systems, or resource management concepts. Strong understanding of distributed systems concepts such as coordination, consistency, and fault tolerance. Experience working with PostgreSQL. Experience designing APIs and backend services. Familiarity with containerised environments and microservices architectures. Strong problem solving and analytical skills. Preferred Experience with Slurm or similar workload managers. Experience in HPC or large scale compute environments. Experience with Golang or Java. Familiarity with C/C++ and performance critical systems. Experience providing technical or project leadership. Benefits Competitive salary commensurate with experience. Highly attractive bonus scheme. Initial 22 days annual leave with future increases, complemented by a flexible buying and selling holiday program. Company pension with generous employer contribution. Wellbeing Unmind app - puts you in control of your mental health. A flexible benefits platform with numerous discount schemes - gym membership, restaurants, cinema tickets, and much more. Regular social club events, spontaneous reward events throughout the year. Cycle purchase scheme. Flexible Private Medical & Dental care programmes. Sponsorship of visas/comprehensive relocation packages. Bank Holiday Swap - our holiday swap program allows you to change it for another day of your choice. Relaxed dress code policy. Learning and Development: We foster a culture of continuous learning and provide tailored training programs through our Learning Hub, designed to enhance technical, commercial, and personal growth. We Care About The Environment: We encourage and actively support a strong sense of community, through volunteering and various company initiatives, as well as a strong company commitment to protecting our environment through sustainable solutions, energy saving and waste reduction enterprises. We see things differently. Diversity fuels our innovation, we value the unique ways in which we differ, and we are committed to equal employment opportunities for all professionals.
29/06/2026
Full time
Software Developer - Distributed Scheduling & Workload Orchestration This role focuses on building scalable and reliable systems that coordinate workloads across clusters, using technologies such as Slurm, Golang, Java, PostgreSQL, and containerised microservices. Responsibilities Design and develop systems for job scheduling and workload orchestration; integrate and extend scheduling capabilities using tools such as Slurm; manage job lifecycles, resource allocation, and execution workflows. Design and build APIs and backend services supporting scheduling systems; work with PostgreSQL to manage system state and coordination. Analyse and improve system performance, scalability, and reliability; ensure efficient resource utilisation across distributed environments. Participate in system design and architecture discussions; collaborate with cross functional teams to evolve scheduling and orchestration capabilities. Qualifications Strong software development experience. Proven experience building backend services or distributed systems. Experience with job scheduling, orchestration systems, or resource management concepts. Strong understanding of distributed systems concepts such as coordination, consistency, and fault tolerance. Experience working with PostgreSQL. Experience designing APIs and backend services. Familiarity with containerised environments and microservices architectures. Strong problem solving and analytical skills. Preferred Experience with Slurm or similar workload managers. Experience in HPC or large scale compute environments. Experience with Golang or Java. Familiarity with C/C++ and performance critical systems. Experience providing technical or project leadership. Benefits Competitive salary commensurate with experience. Highly attractive bonus scheme. Initial 22 days annual leave with future increases, complemented by a flexible buying and selling holiday program. Company pension with generous employer contribution. Wellbeing Unmind app - puts you in control of your mental health. A flexible benefits platform with numerous discount schemes - gym membership, restaurants, cinema tickets, and much more. Regular social club events, spontaneous reward events throughout the year. Cycle purchase scheme. Flexible Private Medical & Dental care programmes. Sponsorship of visas/comprehensive relocation packages. Bank Holiday Swap - our holiday swap program allows you to change it for another day of your choice. Relaxed dress code policy. Learning and Development: We foster a culture of continuous learning and provide tailored training programs through our Learning Hub, designed to enhance technical, commercial, and personal growth. We Care About The Environment: We encourage and actively support a strong sense of community, through volunteering and various company initiatives, as well as a strong company commitment to protecting our environment through sustainable solutions, energy saving and waste reduction enterprises. We see things differently. Diversity fuels our innovation, we value the unique ways in which we differ, and we are committed to equal employment opportunities for all professionals.
Distributed Scheduling & Orchestration Engineer
CGG Services (UK) Limited Crawley, Sussex
CGG Services (UK) Limited in Crawley is looking for a Software Developer specializing in Distributed Scheduling and Workload Orchestration. The role involves building scalable systems using technologies like Slurm, Golang, and PostgreSQL. Responsibilities include designing job scheduling systems and collaborating with cross-functional teams. Successful candidates should have strong software development experience and an understanding of distributed systems. The job offers a competitive salary and a range of benefits, including a flexible holiday program and well-being resources.
29/06/2026
Full time
CGG Services (UK) Limited in Crawley is looking for a Software Developer specializing in Distributed Scheduling and Workload Orchestration. The role involves building scalable systems using technologies like Slurm, Golang, and PostgreSQL. Responsibilities include designing job scheduling systems and collaborating with cross-functional teams. Successful candidates should have strong software development experience and an understanding of distributed systems. The job offers a competitive salary and a range of benefits, including a flexible holiday program and well-being resources.

Modal Window

  • Home
  • Contact
  • About Us
  • FAQs
  • Terms & Conditions
  • Privacy
  • Employer
  • Post a Job
  • Search Resumes
  • Sign in
  • Job Seeker
  • Find Jobs
  • Create Resume
  • Sign in
  • IT blog
  • Facebook
  • Twitter
  • LinkedIn
  • Youtube
© 2008-2026 IT Job Board