Perplexity is revolutionizing how people discover and interact with information through AI-powered search and knowledge tools. As we expand our global footprint, we're establishing a strategic presence in London to drive innovation and growth across Europe. The Role We're seeking an exceptional Site Lead to establish and scale our London office. This is a unique opportunity to shape Perplexity's presence in one of the world's leading tech hubs, building teams and culture from the ground up while driving technical excellence in infrastructure and AI systems. As Site Lead, you'll serve as the face of Perplexity in London, responsible for building our technical organization, fostering a world class engineering culture, and directly managing one or more infrastructure teams. You'll report to senior leadership and work cross functionally with teams across our global footprint. The individual in this role will manage teams in LON themselves while also facilitating. Responsibilities Site Leadership & Culture Establish and lead Perplexity's London office, setting the cultural foundation and operating principles Build a collaborative, high performance engineering culture that aligns with Perplexity's values while embracing the strengths of the London tech ecosystem Serve as the primary point of contact for all London based activities and represent the site in company wide strategic discussions Partner with People/HR, Finance, and Operations to ensure seamless site operations Drive local community engagement, partnerships, and Perplexity's brand presence in the London and European tech community Technical Leadership Directly manage and mentor one or more infrastructure or AI infrastructure teams in London (5-15+ engineers) Set technical direction and strategy for London-based infrastructure initiatives in alignment with company-wide goals Drive architectural decisions and technical excellence across teams Ensure robust systems for deployment, monitoring, scalability, and reliability of infrastructure supporting AI/ML workloads Collaborate with engineering leaders globally to align on technical standards, best practices, and cross site initiatives Team Building & Talent Build and scale high performing infrastructure and AI infrastructure teams through strategic hiring Develop and execute talent acquisition strategy for the London site in partnership with recruiting Create career development frameworks and growth opportunities for engineers Foster technical mentorship and knowledge sharing across teams and sites Cross functional Collaboration Partner with Product, Engineering, and Research teams globally to understand infrastructure needs and deliver solutions Coordinate with other site leads and engineering leaders to ensure effective cross site collaboration Contribute to company wide infrastructure strategy and roadmap planning Facilitate knowledge transfer and best practice sharing across global teams Qualifications Required 10+ years of experience in software engineering with 5+ years in infrastructure, cloud infrastructure, or AI infrastructure roles 3+ years of people management experience, including building and scaling teams Proven track record of establishing or significantly growing an engineering site or office Deep technical expertise in distributed systems, cloud platforms (AWS, GCP, or Azure), and infrastructure automation Experience with infrastructure supporting large scale AI/ML systems, including: GPU infrastructure and orchestration ML training and inference pipelines Model serving and deployment at scale Strong understanding of modern infrastructure technologies: Kubernetes, Terraform, container orchestration, CI/CD systems Demonstrated ability to set technical vision and drive execution across multiple teams Excellent communication and stakeholder management skills Experience working in fast paced, high growth technology companies Passion for building inclusive, diverse, and high performing teams Preferred Experience at companies focused on AI/ML, search, or large scale consumer applications Previous experience as a site lead, office lead, or similar multi team leadership role Background in building infrastructure for LLM training or inference Contributions to open source infrastructure or AI infrastructure projects Experience scaling teams from 0 to 20+ engineers Active involvement in the London or European tech community MBA or advanced technical degree What Success Looks Like 30 Days Deep understanding of Perplexity's infrastructure, technology stack, and organizational structure Established relationships with key stakeholders across engineering, product, and leadership Initial hiring plan and culture strategy for London site established Help the Search, API, AI and Infra teams build out their hiring pipelines 90 Days Core infrastructure team established and ramping in London Clear technical roadmap and priorities defined for London based teams Site culture and operating rhythms established (team meetings, all hands, cross site syncs) London office actively participating in company wide infrastructure initiatives 1 Year London site operating as a high functioning hub with 15-30+ engineers Infrastructure teams delivering measurable impact on system reliability, performance, and scalability Strong talent brand established in London market with healthy hiring pipeline London recognized internally as a strategic site contributing to Perplexity's technical leadership Why Join Perplexity Ground floor opportunity to build and lead a strategic site for a fast growing AI company Work on cutting edge AI infrastructure challenges at massive scale Shape the culture and technical direction of an entire office Competitive compensation including equity Comprehensive benefits package Flexible work environment Opportunity to make a significant impact on how millions of people access and interact with information Location: London, United Kingdom (Hybrid) Perplexity is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
22/07/2026
Full time
Perplexity is revolutionizing how people discover and interact with information through AI-powered search and knowledge tools. As we expand our global footprint, we're establishing a strategic presence in London to drive innovation and growth across Europe. The Role We're seeking an exceptional Site Lead to establish and scale our London office. This is a unique opportunity to shape Perplexity's presence in one of the world's leading tech hubs, building teams and culture from the ground up while driving technical excellence in infrastructure and AI systems. As Site Lead, you'll serve as the face of Perplexity in London, responsible for building our technical organization, fostering a world class engineering culture, and directly managing one or more infrastructure teams. You'll report to senior leadership and work cross functionally with teams across our global footprint. The individual in this role will manage teams in LON themselves while also facilitating. Responsibilities Site Leadership & Culture Establish and lead Perplexity's London office, setting the cultural foundation and operating principles Build a collaborative, high performance engineering culture that aligns with Perplexity's values while embracing the strengths of the London tech ecosystem Serve as the primary point of contact for all London based activities and represent the site in company wide strategic discussions Partner with People/HR, Finance, and Operations to ensure seamless site operations Drive local community engagement, partnerships, and Perplexity's brand presence in the London and European tech community Technical Leadership Directly manage and mentor one or more infrastructure or AI infrastructure teams in London (5-15+ engineers) Set technical direction and strategy for London-based infrastructure initiatives in alignment with company-wide goals Drive architectural decisions and technical excellence across teams Ensure robust systems for deployment, monitoring, scalability, and reliability of infrastructure supporting AI/ML workloads Collaborate with engineering leaders globally to align on technical standards, best practices, and cross site initiatives Team Building & Talent Build and scale high performing infrastructure and AI infrastructure teams through strategic hiring Develop and execute talent acquisition strategy for the London site in partnership with recruiting Create career development frameworks and growth opportunities for engineers Foster technical mentorship and knowledge sharing across teams and sites Cross functional Collaboration Partner with Product, Engineering, and Research teams globally to understand infrastructure needs and deliver solutions Coordinate with other site leads and engineering leaders to ensure effective cross site collaboration Contribute to company wide infrastructure strategy and roadmap planning Facilitate knowledge transfer and best practice sharing across global teams Qualifications Required 10+ years of experience in software engineering with 5+ years in infrastructure, cloud infrastructure, or AI infrastructure roles 3+ years of people management experience, including building and scaling teams Proven track record of establishing or significantly growing an engineering site or office Deep technical expertise in distributed systems, cloud platforms (AWS, GCP, or Azure), and infrastructure automation Experience with infrastructure supporting large scale AI/ML systems, including: GPU infrastructure and orchestration ML training and inference pipelines Model serving and deployment at scale Strong understanding of modern infrastructure technologies: Kubernetes, Terraform, container orchestration, CI/CD systems Demonstrated ability to set technical vision and drive execution across multiple teams Excellent communication and stakeholder management skills Experience working in fast paced, high growth technology companies Passion for building inclusive, diverse, and high performing teams Preferred Experience at companies focused on AI/ML, search, or large scale consumer applications Previous experience as a site lead, office lead, or similar multi team leadership role Background in building infrastructure for LLM training or inference Contributions to open source infrastructure or AI infrastructure projects Experience scaling teams from 0 to 20+ engineers Active involvement in the London or European tech community MBA or advanced technical degree What Success Looks Like 30 Days Deep understanding of Perplexity's infrastructure, technology stack, and organizational structure Established relationships with key stakeholders across engineering, product, and leadership Initial hiring plan and culture strategy for London site established Help the Search, API, AI and Infra teams build out their hiring pipelines 90 Days Core infrastructure team established and ramping in London Clear technical roadmap and priorities defined for London based teams Site culture and operating rhythms established (team meetings, all hands, cross site syncs) London office actively participating in company wide infrastructure initiatives 1 Year London site operating as a high functioning hub with 15-30+ engineers Infrastructure teams delivering measurable impact on system reliability, performance, and scalability Strong talent brand established in London market with healthy hiring pipeline London recognized internally as a strategic site contributing to Perplexity's technical leadership Why Join Perplexity Ground floor opportunity to build and lead a strategic site for a fast growing AI company Work on cutting edge AI infrastructure challenges at massive scale Shape the culture and technical direction of an entire office Competitive compensation including equity Comprehensive benefits package Flexible work environment Opportunity to make a significant impact on how millions of people access and interact with information Location: London, United Kingdom (Hybrid) Perplexity is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
the company is revolutionizing how people discover and interact with information through AI-powered search and knowledge tools. As we expand our global footprint, we're establishing a strategic presence in London to drive innovation and growth across Europe. The Role: We're seeking an exceptional Site Lead to establish and scale our London office. This is a unique opportunity to shape the company's presence in one of the world's leading tech hubs, building teams and culture from the ground up while driving technical excellence in infrastructure and AI systems. As Site Lead, you'll serve as the face of the company in London, responsible for building our technical organization, fostering a world-class engineering culture, and directly managing one or more infrastructure teams. You'll report to senior leadership and work cross-functionally with teams across our global footprint. The individual in this role will manage teams in LON themselves while also facilitating. Responsibilities: Site Leadership & Culture Establish and lead the company's London office, setting the cultural foundation and operating principles Build a collaborative, high-performance engineering culture that aligns with the company's values while embracing the strengths of the London tech ecosystem Serve as the primary point of contact for all London-based activities and represent the site in company-wide strategic discussions Partner with People/HR, Finance, and Operations to ensure seamless site operations Drive local community engagement, partnerships, and the company's brand presence in the London and European tech community Technical Leadership: Directly manage and mentor one or more infrastructure or AI infrastructure teams in London (5-15+ engineers) Set technical direction and strategy for London-based infrastructure initiatives in alignment with company-wide goals Drive architectural decisions and technical excellence across teams Ensure robust systems for deployment, monitoring, scalability, and reliability of infrastructure supporting AI/ML workloads Collaborate with engineering leaders globally to align on technical standards, best practices, and cross-site initiatives Team Building & Talent: Build and scale high-performing infrastructure and AI infrastructure teams through strategic hiring Develop and execute talent acquisition strategy for the London site in partnership with recruiting Create career development frameworks and growth opportunities for engineers Foster technical mentorship and knowledge sharing across teams and sites Cross-functional Collaboration: Partner with Product, Engineering, and Research teams globally to understand infrastructure needs and deliver solutions Coordinate with other site leads and engineering leaders to ensure effective cross-site collaboration Contribute to company-wide infrastructure strategy and roadmap planning Facilitate knowledge transfer and best practice sharing across global teams Qualifications: Required 10+ years of experience in software engineering with 5+ years in infrastructure, cloud infrastructure, or AI infrastructure roles 3+ years of people management experience, including building and scaling teams Proven track record of establishing or significantly growing an engineering site or office Deep technical expertise in distributed systems, cloud platforms (AWS, GCP, or Azure), and infrastructure automation Experience with infrastructure supporting large-scale AI/ML systems, including: GPU infrastructure and orchestration ML training and inference pipelines Model serving and deployment at scale Strong understanding of modern infrastructure technologies: Kubernetes, Terraform, container orchestration, CI/CD systems Demonstrated ability to set technical vision and drive execution across multiple teams Excellent communication and stakeholder management skills Experience working in fast-paced, high-growth technology companies Passion for building inclusive, diverse, and high-performing teams Preferred Experience at companies focused on AI/ML, search, or large-scale consumer applications Previous experience as a site lead, office lead, or similar multi-team leadership role Background in building infrastructure for LLM training or inference Contributions to open-source infrastructure or AI infrastructure projects Experience scaling teams from 0 to 20+ engineers Active involvement in the London or European tech community MBA or advanced technical degree What Success Looks Like: 30 Days Deep understanding of the company's infrastructure, technology stack, and organizational structure Established relationships with key stakeholders across engineering, product, and leadership Initial hiring plan and culture strategy for London site established Help the Search, API, AI and Infra teams build out their hiring pipelines 90 Days Core infrastructure team established and ramping in London Clear technical roadmap and priorities defined for London-based teams Site culture and operating rhythms established (team meetings, all-hands, cross-site syncs) London office actively participating in company-wide infrastructure initiatives 1 Year London site operating as a high-functioning hub with 15-30+ engineers Infrastructure teams delivering measurable impact on system reliability, performance, and scalability Strong talent brand established in London market with healthy hiring pipeline London recognized internally as a strategic site contributing to the company's technical leadership Why Join the company Ground-floor opportunity to build and lead a strategic site for a fast-growing AI company Work on cutting-edge AI infrastructure challenges at massive scale Shape the culture and technical direction of an entire office Competitive compensation including equity Comprehensive benefits package Flexible work environment Opportunity to make a significant impact on how millions of people access and interact with information Location: London, United Kingdom (Hybrid) the company is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
15/07/2026
Full time
the company is revolutionizing how people discover and interact with information through AI-powered search and knowledge tools. As we expand our global footprint, we're establishing a strategic presence in London to drive innovation and growth across Europe. The Role: We're seeking an exceptional Site Lead to establish and scale our London office. This is a unique opportunity to shape the company's presence in one of the world's leading tech hubs, building teams and culture from the ground up while driving technical excellence in infrastructure and AI systems. As Site Lead, you'll serve as the face of the company in London, responsible for building our technical organization, fostering a world-class engineering culture, and directly managing one or more infrastructure teams. You'll report to senior leadership and work cross-functionally with teams across our global footprint. The individual in this role will manage teams in LON themselves while also facilitating. Responsibilities: Site Leadership & Culture Establish and lead the company's London office, setting the cultural foundation and operating principles Build a collaborative, high-performance engineering culture that aligns with the company's values while embracing the strengths of the London tech ecosystem Serve as the primary point of contact for all London-based activities and represent the site in company-wide strategic discussions Partner with People/HR, Finance, and Operations to ensure seamless site operations Drive local community engagement, partnerships, and the company's brand presence in the London and European tech community Technical Leadership: Directly manage and mentor one or more infrastructure or AI infrastructure teams in London (5-15+ engineers) Set technical direction and strategy for London-based infrastructure initiatives in alignment with company-wide goals Drive architectural decisions and technical excellence across teams Ensure robust systems for deployment, monitoring, scalability, and reliability of infrastructure supporting AI/ML workloads Collaborate with engineering leaders globally to align on technical standards, best practices, and cross-site initiatives Team Building & Talent: Build and scale high-performing infrastructure and AI infrastructure teams through strategic hiring Develop and execute talent acquisition strategy for the London site in partnership with recruiting Create career development frameworks and growth opportunities for engineers Foster technical mentorship and knowledge sharing across teams and sites Cross-functional Collaboration: Partner with Product, Engineering, and Research teams globally to understand infrastructure needs and deliver solutions Coordinate with other site leads and engineering leaders to ensure effective cross-site collaboration Contribute to company-wide infrastructure strategy and roadmap planning Facilitate knowledge transfer and best practice sharing across global teams Qualifications: Required 10+ years of experience in software engineering with 5+ years in infrastructure, cloud infrastructure, or AI infrastructure roles 3+ years of people management experience, including building and scaling teams Proven track record of establishing or significantly growing an engineering site or office Deep technical expertise in distributed systems, cloud platforms (AWS, GCP, or Azure), and infrastructure automation Experience with infrastructure supporting large-scale AI/ML systems, including: GPU infrastructure and orchestration ML training and inference pipelines Model serving and deployment at scale Strong understanding of modern infrastructure technologies: Kubernetes, Terraform, container orchestration, CI/CD systems Demonstrated ability to set technical vision and drive execution across multiple teams Excellent communication and stakeholder management skills Experience working in fast-paced, high-growth technology companies Passion for building inclusive, diverse, and high-performing teams Preferred Experience at companies focused on AI/ML, search, or large-scale consumer applications Previous experience as a site lead, office lead, or similar multi-team leadership role Background in building infrastructure for LLM training or inference Contributions to open-source infrastructure or AI infrastructure projects Experience scaling teams from 0 to 20+ engineers Active involvement in the London or European tech community MBA or advanced technical degree What Success Looks Like: 30 Days Deep understanding of the company's infrastructure, technology stack, and organizational structure Established relationships with key stakeholders across engineering, product, and leadership Initial hiring plan and culture strategy for London site established Help the Search, API, AI and Infra teams build out their hiring pipelines 90 Days Core infrastructure team established and ramping in London Clear technical roadmap and priorities defined for London-based teams Site culture and operating rhythms established (team meetings, all-hands, cross-site syncs) London office actively participating in company-wide infrastructure initiatives 1 Year London site operating as a high-functioning hub with 15-30+ engineers Infrastructure teams delivering measurable impact on system reliability, performance, and scalability Strong talent brand established in London market with healthy hiring pipeline London recognized internally as a strategic site contributing to the company's technical leadership Why Join the company Ground-floor opportunity to build and lead a strategic site for a fast-growing AI company Work on cutting-edge AI infrastructure challenges at massive scale Shape the culture and technical direction of an entire office Competitive compensation including equity Comprehensive benefits package Flexible work environment Opportunity to make a significant impact on how millions of people access and interact with information Location: London, United Kingdom (Hybrid) the company is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
Intercom is the AI Customer Service company on a mission to help businesses provide incredible customer experiences. Our AI agent Fin, the most advanced customer service AI agent on the market, lets businesses deliver always on, impeccable customer service and ultimately transform their customer experiences for the better. Fin can also be combined with our Helpdesk to become a complete solution called the Intercom Customer Service Suite, which provides AI enhanced support for the more complex or high touch queries that require a human agent. Founded in 2011 and trusted by nearly 30,000 global businesses, Intercom is setting the new standard for customer service. Driven by our core values, we push boundaries, build with speed and intensity, and consistently deliver incredible value to our customers. What's the opportunity? We're looking for Senior+ AI Infrastructure Engineers to build the systems that train and serve Intercom's next generation of AI products. Intercom is an AI company that builds from the GPU all the way up to a user agent that resolves millions of customer service queries a month. You'll join a small, highly technical team working at the cutting edge of modern AI infrastructure. The AI Infra team built the training pipelines and runs the inference for custom models like Fin Apex, which outperforms frontier models in customer service tasks, and is the foundation of the AI Group's full stack approach to AI. We're particularly interested in engineers who have: A track record of working on model training or model inference at scale, or on low level GPU coding (e.g. CUDA, Triton). Experience with one is great, multiple is even better. What will I be doing? As a Senior AI Infrastructure Engineer focused on model training and inference, you will: Implement and scale training pipelines for large transformer and LLM models, from data ingestion and preprocessing through distributed training and evaluation. Build and optimize inference services that deliver low latency, high reliability experiences for our customers, including autoscaling, routing, and fallbacks. Work on GPU level performance: tuning kernels, improving utilization, and identifying bottlenecks across our training and inference stack. Collaborate closely with ML scientists to implement cutting edge training and inference methods and bring them to production. Play an active role in hiring, mentoring, and developing other engineers on the team. Raise the bar for technical standards, reliability, and operational excellence across Intercom's AI platform. These are indicative, not hard requirements. We're looking to hire Senior+ AI Infrastructure Engineers. You're likely a great fit if: You have 5+ years of experience in software engineering, with a strong track record of shipping high quality products or platforms. You hold a degree in Computer Science, Computer Engineering, or a related field (or you have equivalent experience with very strong fundamentals). You have hands on experience with one or more of the following: Model training (especially transformers and LLMs). Model inference at scale (again, especially transformers and LLMs). Low level GPU work, such as writing CUDA or Triton kernels. Comfortable working in production environments at meaningful scale (traffic, data, or organizational). You communicate clearly, can explain complex technical topics to different audiences, and enjoy close collaboration with both engineers and non engineers. You take pride in strong technical fundamentals, love learning, and are willing to invest in your own development. Have deep knowledge of at least one programming language (for example Python, Ruby, Java, Go, etc.). Specific language experience is less important than your ability to write clean, reliable code and learn new stacks quickly. None of these are required, but they're nice to have: Experience at AI native companies that train and/or run inference for their own models (e.g. modern AI labs or AI native product companies). Experience running training or inference workloads on Kubernetes. Experience with AWS or other major cloud providers. Production experience with Python in ML or infrastructure contexts. Demonstrated passion for technology through personal projects, open source, meetups, or publishing content about your work and learnings. Benefits Competitive salary and equity in a fast-growing start up. We serve lunch every weekday, plus a variety of snack foods and a fully stocked kitchen. Unlimited access to Claude Code and best class AI tools; experimentation & building is encouraged & celebrated. Pension scheme & match up to 4%. Peace of mind with life assurance, as well as comprehensive health and dental insurance for you and your dependents. Flexible paid time off policy. Paid maternity leave, as well as 6 weeks paternity leave for fathers, to let you spend valuable time with your loved ones. If you're cycling, we've got you covered on the Cycle to Work Scheme. With secure bike storage too. MacBooks are our standard, but we also offer Windows for certain roles when needed. Intercom has a hybrid working policy. We believe that working in person helps us stay connected, collaborate easier and create a great culture while still providing flexibility to work from home. We expect employees to be in the office at least three days per week. We have a radically open and accepting culture at Intercom. We avoid spending time on divisive subjects to foster a safe and cohesive work environment for everyone. As an organization, our policy is to not advocate on behalf of the company or our employees on any social or political topics out of our internal or external communications. We respect personal opinion and expression on these topics on personal social platforms on personal time, and do not challenge or confront anyone for their views on non work related topics. Our goal is to focus on doing incredible work to achieve our goals and unite the company through our core values. Intercom values diversity and is committed to a policy of Equal Employment Opportunity. Intercom will not discriminate against an applicant or employee on the basis of race, color, religion, creed, national origin, ancestry, sex, gender, age, physical or mental disability, veteran or military status, genetic information, sexual orientation, gender identity, gender expression, marital status, or any other legally recognized protected basis under federal, state, or local law.
14/07/2026
Full time
Intercom is the AI Customer Service company on a mission to help businesses provide incredible customer experiences. Our AI agent Fin, the most advanced customer service AI agent on the market, lets businesses deliver always on, impeccable customer service and ultimately transform their customer experiences for the better. Fin can also be combined with our Helpdesk to become a complete solution called the Intercom Customer Service Suite, which provides AI enhanced support for the more complex or high touch queries that require a human agent. Founded in 2011 and trusted by nearly 30,000 global businesses, Intercom is setting the new standard for customer service. Driven by our core values, we push boundaries, build with speed and intensity, and consistently deliver incredible value to our customers. What's the opportunity? We're looking for Senior+ AI Infrastructure Engineers to build the systems that train and serve Intercom's next generation of AI products. Intercom is an AI company that builds from the GPU all the way up to a user agent that resolves millions of customer service queries a month. You'll join a small, highly technical team working at the cutting edge of modern AI infrastructure. The AI Infra team built the training pipelines and runs the inference for custom models like Fin Apex, which outperforms frontier models in customer service tasks, and is the foundation of the AI Group's full stack approach to AI. We're particularly interested in engineers who have: A track record of working on model training or model inference at scale, or on low level GPU coding (e.g. CUDA, Triton). Experience with one is great, multiple is even better. What will I be doing? As a Senior AI Infrastructure Engineer focused on model training and inference, you will: Implement and scale training pipelines for large transformer and LLM models, from data ingestion and preprocessing through distributed training and evaluation. Build and optimize inference services that deliver low latency, high reliability experiences for our customers, including autoscaling, routing, and fallbacks. Work on GPU level performance: tuning kernels, improving utilization, and identifying bottlenecks across our training and inference stack. Collaborate closely with ML scientists to implement cutting edge training and inference methods and bring them to production. Play an active role in hiring, mentoring, and developing other engineers on the team. Raise the bar for technical standards, reliability, and operational excellence across Intercom's AI platform. These are indicative, not hard requirements. We're looking to hire Senior+ AI Infrastructure Engineers. You're likely a great fit if: You have 5+ years of experience in software engineering, with a strong track record of shipping high quality products or platforms. You hold a degree in Computer Science, Computer Engineering, or a related field (or you have equivalent experience with very strong fundamentals). You have hands on experience with one or more of the following: Model training (especially transformers and LLMs). Model inference at scale (again, especially transformers and LLMs). Low level GPU work, such as writing CUDA or Triton kernels. Comfortable working in production environments at meaningful scale (traffic, data, or organizational). You communicate clearly, can explain complex technical topics to different audiences, and enjoy close collaboration with both engineers and non engineers. You take pride in strong technical fundamentals, love learning, and are willing to invest in your own development. Have deep knowledge of at least one programming language (for example Python, Ruby, Java, Go, etc.). Specific language experience is less important than your ability to write clean, reliable code and learn new stacks quickly. None of these are required, but they're nice to have: Experience at AI native companies that train and/or run inference for their own models (e.g. modern AI labs or AI native product companies). Experience running training or inference workloads on Kubernetes. Experience with AWS or other major cloud providers. Production experience with Python in ML or infrastructure contexts. Demonstrated passion for technology through personal projects, open source, meetups, or publishing content about your work and learnings. Benefits Competitive salary and equity in a fast-growing start up. We serve lunch every weekday, plus a variety of snack foods and a fully stocked kitchen. Unlimited access to Claude Code and best class AI tools; experimentation & building is encouraged & celebrated. Pension scheme & match up to 4%. Peace of mind with life assurance, as well as comprehensive health and dental insurance for you and your dependents. Flexible paid time off policy. Paid maternity leave, as well as 6 weeks paternity leave for fathers, to let you spend valuable time with your loved ones. If you're cycling, we've got you covered on the Cycle to Work Scheme. With secure bike storage too. MacBooks are our standard, but we also offer Windows for certain roles when needed. Intercom has a hybrid working policy. We believe that working in person helps us stay connected, collaborate easier and create a great culture while still providing flexibility to work from home. We expect employees to be in the office at least three days per week. We have a radically open and accepting culture at Intercom. We avoid spending time on divisive subjects to foster a safe and cohesive work environment for everyone. As an organization, our policy is to not advocate on behalf of the company or our employees on any social or political topics out of our internal or external communications. We respect personal opinion and expression on these topics on personal social platforms on personal time, and do not challenge or confront anyone for their views on non work related topics. Our goal is to focus on doing incredible work to achieve our goals and unite the company through our core values. Intercom values diversity and is committed to a policy of Equal Employment Opportunity. Intercom will not discriminate against an applicant or employee on the basis of race, color, religion, creed, national origin, ancestry, sex, gender, age, physical or mental disability, veteran or military status, genetic information, sexual orientation, gender identity, gender expression, marital status, or any other legally recognized protected basis under federal, state, or local law.
What We're Looking For We are seeking a DevOps Engineer to build and own the infrastructure that underpins our AI driven materials discovery platform. You'll work directly with world renowned ML researchers and software engineers to accelerate real scientific breakthroughs by making model training, experimentation, and deployment fast, reliable, and reproducible. This is a foundational hire. You'll set the patterns others build on. You will be joining a small, highly ambitious team of world renowned engineers, AI researchers, and materials scientists. We move fast and value people who are energised by that. What You'll Do Design, provision, and manage cloud infrastructure (AWS/GCP) using infrastructure as code; Terraform, Pulumi, or equivalent. Own GPU compute environments for model training and inference, including cluster configuration, job scheduling, and cost optimisation. Build and maintain CI/CD pipelines that support rapid model iteration, automated testing, and safe deployments. Support ML workflow orchestration; experiment tracking, training run management, and data pipeline reliability. Ensure reproducibility across research and production environments through containerisation and rigorous environment management. Define monitoring, alerting, and incident response processes so the team can move fast without things silently breaking. Implement security best practices: secrets management, IAM, network segmentation, vulnerability scanning. Build internal tooling and documentation that lets researchers self serve infrastructure without waiting on you. Skills & Qualifications 4+ years in a DevOps, Platform Engineering, or SRE role. Strong proficiency with at least one major cloud provider and its core services (compute, storage, networking, IAM). Hands on experience with infrastructure as code and container orchestration (Kubernetes or equivalent). Solid CI/CD pipeline experience, GitHub Actions, GitLab CI, or similar. Proficient in Python and Bash; comfortable reading and writing code across a polyglot stack. Deep Linux systems knowledge and strong networking fundamentals. A bias for building things properly the first time, even under early stage constraints. Nice to Have Experience with GPU cluster management and ML training workloads (NVIDIA, CUDA, distributed training). Familiarity with MLOps tooling: Experiment tracking (MLflow, Weights & Biases). Workflow orchestration (Airflow, Prefect, Argo). Data versioning (DVC). Background in scientific computing or HPC environments. Prior experience at a deep tech or computational science company. Why Join Us Work directly on infrastructure that enables AI to make real scientific discoveries. Shape how we build from day one, no legacy systems, no inherited mess. Collaborate with world class researchers across materials science and machine learning. Diffractive is building the AI Material Scientist that autonomously learns from real world experimentation to push the boundaries of scientific discovery. We're early, moving fast, and working on problems that genuinely matter. We are a London based company with a flexible approach to how and where you work. We offer competitive salary, generous equity, and benefits. You'll have a real stake in what you build and in the company's overall success. Equal Opportunity Diffractive is an equal opportunities employer. We are committed to creating an inclusive environment for all employees and welcome applications from people of all backgrounds, experiences, and identities. If you require any adjustments or accommodations at any point during the interview process please let us know - we will be happy to help.
11/07/2026
Full time
What We're Looking For We are seeking a DevOps Engineer to build and own the infrastructure that underpins our AI driven materials discovery platform. You'll work directly with world renowned ML researchers and software engineers to accelerate real scientific breakthroughs by making model training, experimentation, and deployment fast, reliable, and reproducible. This is a foundational hire. You'll set the patterns others build on. You will be joining a small, highly ambitious team of world renowned engineers, AI researchers, and materials scientists. We move fast and value people who are energised by that. What You'll Do Design, provision, and manage cloud infrastructure (AWS/GCP) using infrastructure as code; Terraform, Pulumi, or equivalent. Own GPU compute environments for model training and inference, including cluster configuration, job scheduling, and cost optimisation. Build and maintain CI/CD pipelines that support rapid model iteration, automated testing, and safe deployments. Support ML workflow orchestration; experiment tracking, training run management, and data pipeline reliability. Ensure reproducibility across research and production environments through containerisation and rigorous environment management. Define monitoring, alerting, and incident response processes so the team can move fast without things silently breaking. Implement security best practices: secrets management, IAM, network segmentation, vulnerability scanning. Build internal tooling and documentation that lets researchers self serve infrastructure without waiting on you. Skills & Qualifications 4+ years in a DevOps, Platform Engineering, or SRE role. Strong proficiency with at least one major cloud provider and its core services (compute, storage, networking, IAM). Hands on experience with infrastructure as code and container orchestration (Kubernetes or equivalent). Solid CI/CD pipeline experience, GitHub Actions, GitLab CI, or similar. Proficient in Python and Bash; comfortable reading and writing code across a polyglot stack. Deep Linux systems knowledge and strong networking fundamentals. A bias for building things properly the first time, even under early stage constraints. Nice to Have Experience with GPU cluster management and ML training workloads (NVIDIA, CUDA, distributed training). Familiarity with MLOps tooling: Experiment tracking (MLflow, Weights & Biases). Workflow orchestration (Airflow, Prefect, Argo). Data versioning (DVC). Background in scientific computing or HPC environments. Prior experience at a deep tech or computational science company. Why Join Us Work directly on infrastructure that enables AI to make real scientific discoveries. Shape how we build from day one, no legacy systems, no inherited mess. Collaborate with world class researchers across materials science and machine learning. Diffractive is building the AI Material Scientist that autonomously learns from real world experimentation to push the boundaries of scientific discovery. We're early, moving fast, and working on problems that genuinely matter. We are a London based company with a flexible approach to how and where you work. We offer competitive salary, generous equity, and benefits. You'll have a real stake in what you build and in the company's overall success. Equal Opportunity Diffractive is an equal opportunities employer. We are committed to creating an inclusive environment for all employees and welcome applications from people of all backgrounds, experiences, and identities. If you require any adjustments or accommodations at any point during the interview process please let us know - we will be happy to help.
About Anthropic Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About The Role Claude has your back. AIRE has Claude's. Help us keep Claude reliable for everyone who depends on it. AIRE (AI Reliability Engineering) partners with teams across Anthropic to improve reliability across our most critical serving paths every hop from the SDK through our network, API layers, serving infrastructure, and accelerators and back. We jump into the trenches alongside partner teams to make the systems that deliver Claude more robust and resilient, be it during an incident or collaborating on projects. Reliability here is an emergent phenomenon that transcends any single team's boundaries, so someone has to zoom out and look at the whole picture. That's us and it means few teams at Anthropic offer this kind of dynamic, cross-cutting exposure to the systems that matter most. Responsibilities Develop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity. Design and implement monitoring and observability systems across the token path. Assist in the design and implementation of high-availability serving infrastructure across multiple regions and cloud providers Lead incident response for critical AI services, ensuring rapid recovery, thorough incident reviews, and systematic improvements. Support the reliability of safeguard model serving critical for both site reliability and Anthropic's safety commitments. You may be a good fit if you Have strong distributed systems, infrastructure, or reliability backgrounds we're looking for reliability-minded software engineers and SREs. Are curious and brave comfortable jumping into unfamiliar systems during an incident and helping drive resolution even when you don't have deep expertise yet. Think holistically about how systems compose and where the seams are. Can build lasting relationships across teams our engagement model depends on being welcomed as teammates, not outsiders with opinions. Care about users and feel ownership over outcomes, even for systems you don't own. Have excellent communication and collaboration skills you'll be partnering across the entire company. Bring diverse experience the team's strength comes from people who've built product stacks, scaled databases, run massive distributed systems, and everything in between. Strong candidates may also Have been an SRE, Production Engineer, or in similar reliability-focused roles on large scale systems Have experience operating large-scale model serving or training infrastructure (>1000 GPUs). Have experience with one or more ML hardware accelerators (GPUs, TPUs, Trainium). Understand ML-specific networking optimizations like RDMA and InfiniBand. Have expertise in AI-specific observability tools and frameworks. Have experience with chaos engineering and systematic resilience testing. Have contributed to open-source infrastructure or ML tooling. The annual compensation range for this role is listed below. For sales roles, the range provided is the role's On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role. Annual Salary £325,000-£390,000 GBP Logistics Education requirements We require at least a Bachelor's degree in a related field or equivalent experience. Location-based hybrid policy Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices. Visa sponsorship We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this. We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team. Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you email addresses. In some cases, we may partner with vetted recruiting agencies who will identify themselves as working on behalf of Anthropic. Be cautious of emails from other domains. Legitimate Anthropic recruiters will never ask for money, fees, or banking information before your first day. If you're ever unsure about a communication, don't click any links-visit directly for confirmed position openings. How We're Different We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact - advancing our long-term goals of steerable, trustworthy AI - rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills. The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences. Come work with us! Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates' AI Usage Learn about our policy for using AI in our application process
09/07/2026
Full time
About Anthropic Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About The Role Claude has your back. AIRE has Claude's. Help us keep Claude reliable for everyone who depends on it. AIRE (AI Reliability Engineering) partners with teams across Anthropic to improve reliability across our most critical serving paths every hop from the SDK through our network, API layers, serving infrastructure, and accelerators and back. We jump into the trenches alongside partner teams to make the systems that deliver Claude more robust and resilient, be it during an incident or collaborating on projects. Reliability here is an emergent phenomenon that transcends any single team's boundaries, so someone has to zoom out and look at the whole picture. That's us and it means few teams at Anthropic offer this kind of dynamic, cross-cutting exposure to the systems that matter most. Responsibilities Develop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity. Design and implement monitoring and observability systems across the token path. Assist in the design and implementation of high-availability serving infrastructure across multiple regions and cloud providers Lead incident response for critical AI services, ensuring rapid recovery, thorough incident reviews, and systematic improvements. Support the reliability of safeguard model serving critical for both site reliability and Anthropic's safety commitments. You may be a good fit if you Have strong distributed systems, infrastructure, or reliability backgrounds we're looking for reliability-minded software engineers and SREs. Are curious and brave comfortable jumping into unfamiliar systems during an incident and helping drive resolution even when you don't have deep expertise yet. Think holistically about how systems compose and where the seams are. Can build lasting relationships across teams our engagement model depends on being welcomed as teammates, not outsiders with opinions. Care about users and feel ownership over outcomes, even for systems you don't own. Have excellent communication and collaboration skills you'll be partnering across the entire company. Bring diverse experience the team's strength comes from people who've built product stacks, scaled databases, run massive distributed systems, and everything in between. Strong candidates may also Have been an SRE, Production Engineer, or in similar reliability-focused roles on large scale systems Have experience operating large-scale model serving or training infrastructure (>1000 GPUs). Have experience with one or more ML hardware accelerators (GPUs, TPUs, Trainium). Understand ML-specific networking optimizations like RDMA and InfiniBand. Have expertise in AI-specific observability tools and frameworks. Have experience with chaos engineering and systematic resilience testing. Have contributed to open-source infrastructure or ML tooling. The annual compensation range for this role is listed below. For sales roles, the range provided is the role's On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role. Annual Salary £325,000-£390,000 GBP Logistics Education requirements We require at least a Bachelor's degree in a related field or equivalent experience. Location-based hybrid policy Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices. Visa sponsorship We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this. We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team. Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you email addresses. In some cases, we may partner with vetted recruiting agencies who will identify themselves as working on behalf of Anthropic. Be cautious of emails from other domains. Legitimate Anthropic recruiters will never ask for money, fees, or banking information before your first day. If you're ever unsure about a communication, don't click any links-visit directly for confirmed position openings. How We're Different We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact - advancing our long-term goals of steerable, trustworthy AI - rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills. The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences. Come work with us! Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates' AI Usage Learn about our policy for using AI in our application process
Senior Machine Learning Infrastructure Engineer London, United Kingdom About us PhysicsX is a deep-tech company with roots in numerical physics and Formula One, dedicated to accelerating hardware innovation at the speed of software. We are building an AI-driven simulation software stack for engineering and manufacturing across advanced industries. By enabling high-fidelity, multi-physics simulation through AI inference across the entire engineering lifecycle, PhysicsX unlocks new levels of optimization and automation in design, manufacturing, and operations - empowering engineers to push the boundaries of possibility. Our customers include leading innovators in Aerospace & Defense, Materials, Energy, Semiconductors, and Automotive. Note:We are currently recruiting for multiple positions, however please only apply for the role that best aligns with your skillset and career goals. The Role The Senior ML Infrastructure Engineer will extend and operate the infrastructure that powers our research model training, fine-tuning, and serving pipelines. You will be embedded within our Research function, partnering directly with ML engineers and research scientists to ensure they can train Large Physics Models efficiently and reliably at scale. Team Context In this role, you will be vertically embedded in Research, working daily with: Research Scientists who determine the model architectures and methods ML Engineers who implement and develop the models Simulation Data Engineers who are accountable for upstream data pipelines You will have end-to-end responsibilities over the research infrastructure, with the autonomy to make architectural decisions and the responsibility to keep data flowing reliably. Horizontally, you will be part of an infrastructure engineering group responsible for infrastructure across the company. What you will do Training Infrastructure Design and operate distributed training infrastructure for neural operator architectures (Transolver, Point Cloud Transformer, etc.) on our large NVIDIA DGX B200 platform. Optimize training pipelines for throughput, fault tolerance, and cost efficiency, including checkpointing strategies, gradient accumulation, and multi-node synchronization. Build and maintain experiment tracking and observability systems that give researchers clear visibility into training runs, hyperparameter sweeps, and model performance. Data I/O and Performance Solve data loading bottlenecks for large-scale mesh datasets. Optimize data pipelines for efficient I/O from cloud storage, including prefetching, caching, and format optimization. Work with heterogeneous data sources of varying formats and resolutions. Model Serving and Deployment Build serving infrastructure for pre-trained LPMs, supporting both zero shot inference and uncertainty quantification (Monte Carlo Dropout). Design and implement model packaging pipelines for customer deployment. Models must run reliably in customer environments with fine tuning capabilities. Ensure reproducibility: any model checkpoint should be deployable with consistent behaviour. Platform and Tooling Improve developer experience for the Research team with fast iteration cycles, reliable CI/CD, clear debugging tools. Collaborate with the broader Infrastructure team on shared patterns and standards. What you bring to the table Ability to scope and effectively deliver projects, prioritising activity as needed. Problem solving skills and the ability to analyse issues, identify causes, and recommend solutions quickly. Excellent collaboration and communication skills, especially in a research setting. You can translate "the model isn't converging" into infrastructure hypotheses and solutions, and can bridge technical abstractions with implementations. 5+ years of experience building and operating ML infrastructure at scale: Deep expertise in distributed training: you've debugged NCCL hangs, optimized collective communication, and know when to use FSDP vs. DDP vs. pipeline parallelism Strong systems fundamentals: Linux, networking (including domain specific NVLink and InfiniBand), storage I/O, profiling and performance optimization Production experience with Kubernetes and SLURM for job orchestration on GPU clusters Proficiency in Python and ML frameworks (PyTorch strongly preferred) Experience with cloud GPU infrastructure; ideally CoreWeave or similar GPU/HPC-focused clouds Ideally Experience with geometric deep learning or neural operators, architectures that operate on meshes, point clouds, or graphs Background in HPC for simulation engineering, familiarity with how CFD/FEA workflows generate and consume data Experience building model serving infrastructure with latency and throughput requirements Familiarity with experiment tracking tools (Weights & Biases, MLflow) and observability stacks (Prometheus, Grafana) What we offer Equity options - share in our success and growth. 10% employer pension contribution - invest in your future. Free office lunches - great food to fuel your workdays. Flexible working - balance your work and life in a way that works for you. Hybrid setup - enjoy our new Shoreditch office while keeping remote flexibility. Enhanced parental leave - support for life's biggest milestones. Private healthcare - comprehensive coverage Personal development - access learning and training to help you grow. Work from anywhere - extend your remote setup to enjoy the sun or reconnect with loved ones. We value diversity and are committed to equal employment opportunity regardless of sex, race, religion, ethnicity, nationality, disability, age, sexual orientation or gender identity. We strongly encourage individuals from groups traditionally underrepresented in tech to apply. To help make a change, we sponsor bright women from disadvantaged backgrounds through their university degrees in science and mathematics. We collect diversity and inclusion data solely for the purpose of monitoring the effectiveness of our equal opportunities policies and ensuring compliance with UK employment and equality legislation. This information is confidential, used only in aggregate form, and will not influence the outcome of your application.
08/07/2026
Full time
Senior Machine Learning Infrastructure Engineer London, United Kingdom About us PhysicsX is a deep-tech company with roots in numerical physics and Formula One, dedicated to accelerating hardware innovation at the speed of software. We are building an AI-driven simulation software stack for engineering and manufacturing across advanced industries. By enabling high-fidelity, multi-physics simulation through AI inference across the entire engineering lifecycle, PhysicsX unlocks new levels of optimization and automation in design, manufacturing, and operations - empowering engineers to push the boundaries of possibility. Our customers include leading innovators in Aerospace & Defense, Materials, Energy, Semiconductors, and Automotive. Note:We are currently recruiting for multiple positions, however please only apply for the role that best aligns with your skillset and career goals. The Role The Senior ML Infrastructure Engineer will extend and operate the infrastructure that powers our research model training, fine-tuning, and serving pipelines. You will be embedded within our Research function, partnering directly with ML engineers and research scientists to ensure they can train Large Physics Models efficiently and reliably at scale. Team Context In this role, you will be vertically embedded in Research, working daily with: Research Scientists who determine the model architectures and methods ML Engineers who implement and develop the models Simulation Data Engineers who are accountable for upstream data pipelines You will have end-to-end responsibilities over the research infrastructure, with the autonomy to make architectural decisions and the responsibility to keep data flowing reliably. Horizontally, you will be part of an infrastructure engineering group responsible for infrastructure across the company. What you will do Training Infrastructure Design and operate distributed training infrastructure for neural operator architectures (Transolver, Point Cloud Transformer, etc.) on our large NVIDIA DGX B200 platform. Optimize training pipelines for throughput, fault tolerance, and cost efficiency, including checkpointing strategies, gradient accumulation, and multi-node synchronization. Build and maintain experiment tracking and observability systems that give researchers clear visibility into training runs, hyperparameter sweeps, and model performance. Data I/O and Performance Solve data loading bottlenecks for large-scale mesh datasets. Optimize data pipelines for efficient I/O from cloud storage, including prefetching, caching, and format optimization. Work with heterogeneous data sources of varying formats and resolutions. Model Serving and Deployment Build serving infrastructure for pre-trained LPMs, supporting both zero shot inference and uncertainty quantification (Monte Carlo Dropout). Design and implement model packaging pipelines for customer deployment. Models must run reliably in customer environments with fine tuning capabilities. Ensure reproducibility: any model checkpoint should be deployable with consistent behaviour. Platform and Tooling Improve developer experience for the Research team with fast iteration cycles, reliable CI/CD, clear debugging tools. Collaborate with the broader Infrastructure team on shared patterns and standards. What you bring to the table Ability to scope and effectively deliver projects, prioritising activity as needed. Problem solving skills and the ability to analyse issues, identify causes, and recommend solutions quickly. Excellent collaboration and communication skills, especially in a research setting. You can translate "the model isn't converging" into infrastructure hypotheses and solutions, and can bridge technical abstractions with implementations. 5+ years of experience building and operating ML infrastructure at scale: Deep expertise in distributed training: you've debugged NCCL hangs, optimized collective communication, and know when to use FSDP vs. DDP vs. pipeline parallelism Strong systems fundamentals: Linux, networking (including domain specific NVLink and InfiniBand), storage I/O, profiling and performance optimization Production experience with Kubernetes and SLURM for job orchestration on GPU clusters Proficiency in Python and ML frameworks (PyTorch strongly preferred) Experience with cloud GPU infrastructure; ideally CoreWeave or similar GPU/HPC-focused clouds Ideally Experience with geometric deep learning or neural operators, architectures that operate on meshes, point clouds, or graphs Background in HPC for simulation engineering, familiarity with how CFD/FEA workflows generate and consume data Experience building model serving infrastructure with latency and throughput requirements Familiarity with experiment tracking tools (Weights & Biases, MLflow) and observability stacks (Prometheus, Grafana) What we offer Equity options - share in our success and growth. 10% employer pension contribution - invest in your future. Free office lunches - great food to fuel your workdays. Flexible working - balance your work and life in a way that works for you. Hybrid setup - enjoy our new Shoreditch office while keeping remote flexibility. Enhanced parental leave - support for life's biggest milestones. Private healthcare - comprehensive coverage Personal development - access learning and training to help you grow. Work from anywhere - extend your remote setup to enjoy the sun or reconnect with loved ones. We value diversity and are committed to equal employment opportunity regardless of sex, race, religion, ethnicity, nationality, disability, age, sexual orientation or gender identity. We strongly encourage individuals from groups traditionally underrepresented in tech to apply. To help make a change, we sponsor bright women from disadvantaged backgrounds through their university degrees in science and mathematics. We collect diversity and inclusion data solely for the purpose of monitoring the effectiveness of our equal opportunities policies and ensuring compliance with UK employment and equality legislation. This information is confidential, used only in aggregate form, and will not influence the outcome of your application.
PhysicsX is a deep-tech company with roots in numerical physics and Formula One, dedicated to accelerating hardware innovation at the speed of software. We are building an AI-driven simulation software stack for engineering and manufacturing across advanced industries. By enabling high-fidelity, multi-physics simulation through AI inference across the entire engineering lifecycle, PhysicsX unlocks new levels of optimization and automation in design, manufacturing, and operations - empowering engineers to push the boundaries of possibility. Our customers include leading innovators in Aerospace & Defense, Materials, Energy, Semiconductors, and Automotive. Note:We are currently recruiting for multiple positions, however please only apply for the role that best aligns with your skillset and career goals. The Role The Senior Simulation Data Engineer will extend and operate the infrastructure that powers our research Data Factory. You will be responsible for the end-to-end pipeline: from geometry preparation and simulation orchestration through validation, post-processing, and delivery to downstream ML training systems, using PhysicsX platform orchestration services where synergies exist. This role sits at the intersection of HPC engineering and data engineering. You will orchestrate long-running CFD simulations at scale, build robust data pipelines, and ensure that every simulation we produce meets rigorous quality standards. Team Context In this role, you will be vertically embedded in Research , working daily with: Research Scientists who define data requirements and quality standards ML Engineers who consume Data Factory outputs for model training ML Infrastructure Engineers who are accountable for downstream training infrastructure You will have end-to-end responsibilities over the Data Factory, with the autonomy to make architectural decisions and the responsibility to keep data flowing reliably. Horizontally, you will be part of an infrastructure engineering group responsible for infrastructure across the company. What you will do Simulation Orchestration Extend and operate the Data Factory infrastructure that orchestrates thousands of CFD simulations per day on cloud compute Design and operate job scheduling systems that maximize throughput while handling failures gracefully Build monitoring and alerting to detect simulation failures, convergence issues, and resource bottlenecks early Build high-performance data pipelines that move simulation outputs from solver results to ML-ready training data Implement geometry preprocessing workflows (mesh preparation, morphing, watertightness validation) Design and operate post-processing pipelines: surface decimation, field interpolation, format conversion Optimize I/O performance for large mesh datasets Data Quality and Validation Implement comprehensive validation checks at every pipeline stage: solver convergence, physical field bounds, post-processing fidelity Build systems that capture and quarantine bad data before they reach training pipelines Track and report data quality metrics across the entire Data Factory Work towards full provenance: training samples should be traceable back to their source geometry and simulation configuration Integration and Delivery Deliver validated datasets to downstream ML training infrastructure in formats optimized for efficient data loading Design data versioning and cataloging systems that support reproducible training runs Work closely with ML Infrastructure Engineers to ensure smooth handoff between data production and model training Support multi-dataset training workflows What you bring to the table Ability to scope and effectively deliver projects, prioritising activity as needed. Problem solving skills and the ability to analyse issues, identify causes, and recommend solutions quickly. Excellent collaboration and communication skills, especially in a research setting. You can translate "the model isn't converging" into infrastructure hypotheses and solutions, and can bridge technical abstractions with implementations. 5+ years of experience in data engineering, HPC engineering, or simulation infrastructure. Strong experience with orchestration systems: SLURM, Kubernetes, Temporal Production data pipeline experience: you've built and operated pipelines that process large volumes of data reliably Proficiency in Python for pipeline development and automation Systems engineering fundamentals: Linux, networking, storage systems, performance debugging Experience with cloud infrastructure; ideally CoreWeave or similar GPU/HPC focused clouds Background in HPC for simulation engineering: experience with CFD, FEA, or similar computational workflows (StarCCM+, OpenFOAM, ANSYS, etc.) Experience with geometry processing: mesh manipulation, CAD formats, PyVista Familiarity with scientific data formats: HDF5, VTK, NetCDF, Zarr Data quality engineering experience: validation frameworks, anomaly detection, data observability Ideally Understanding of CFD fundamentals, enough to interpret solver outputs and validation metrics Experience with 3D geometry pipelines (mesh decimation, field interpolation) Familiarity with ML data loading patterns and how training systems consume data What we offer Equity options - share in our success and growth. 10% employer pension contribution - invest in your future. Free office lunches - great food to fuel your workdays. Flexible working - balance your work and life in a way that works for you. Hybrid setup - enjoy our new Shoreditch office while keeping remote flexibility. Enhanced parental leave - support for life's biggest milestones. Private healthcare - comprehensive coverage Personal development - access learning and training to help you grow. Work from anywhere - extend your remote setup to enjoy the sun or reconnect with loved ones. We value diversity and are committed to equal employment opportunity regardless of sex, race, religion, ethnicity, nationality, disability, age, sexual orientation or gender identity. We strongly encourage individuals from groups traditionally underrepresented in tech to apply. To help make a change, we sponsor bright women from disadvantaged backgrounds through their university degrees in science and mathematics. We collect diversity and inclusion data solely for the purpose of monitoring the effectiveness of our equal opportunities policies and ensuring compliance with UK employment and equality legislation. This information is confidential, used only in aggregate form, and will not influence the outcome of your application.
08/07/2026
Full time
PhysicsX is a deep-tech company with roots in numerical physics and Formula One, dedicated to accelerating hardware innovation at the speed of software. We are building an AI-driven simulation software stack for engineering and manufacturing across advanced industries. By enabling high-fidelity, multi-physics simulation through AI inference across the entire engineering lifecycle, PhysicsX unlocks new levels of optimization and automation in design, manufacturing, and operations - empowering engineers to push the boundaries of possibility. Our customers include leading innovators in Aerospace & Defense, Materials, Energy, Semiconductors, and Automotive. Note:We are currently recruiting for multiple positions, however please only apply for the role that best aligns with your skillset and career goals. The Role The Senior Simulation Data Engineer will extend and operate the infrastructure that powers our research Data Factory. You will be responsible for the end-to-end pipeline: from geometry preparation and simulation orchestration through validation, post-processing, and delivery to downstream ML training systems, using PhysicsX platform orchestration services where synergies exist. This role sits at the intersection of HPC engineering and data engineering. You will orchestrate long-running CFD simulations at scale, build robust data pipelines, and ensure that every simulation we produce meets rigorous quality standards. Team Context In this role, you will be vertically embedded in Research , working daily with: Research Scientists who define data requirements and quality standards ML Engineers who consume Data Factory outputs for model training ML Infrastructure Engineers who are accountable for downstream training infrastructure You will have end-to-end responsibilities over the Data Factory, with the autonomy to make architectural decisions and the responsibility to keep data flowing reliably. Horizontally, you will be part of an infrastructure engineering group responsible for infrastructure across the company. What you will do Simulation Orchestration Extend and operate the Data Factory infrastructure that orchestrates thousands of CFD simulations per day on cloud compute Design and operate job scheduling systems that maximize throughput while handling failures gracefully Build monitoring and alerting to detect simulation failures, convergence issues, and resource bottlenecks early Build high-performance data pipelines that move simulation outputs from solver results to ML-ready training data Implement geometry preprocessing workflows (mesh preparation, morphing, watertightness validation) Design and operate post-processing pipelines: surface decimation, field interpolation, format conversion Optimize I/O performance for large mesh datasets Data Quality and Validation Implement comprehensive validation checks at every pipeline stage: solver convergence, physical field bounds, post-processing fidelity Build systems that capture and quarantine bad data before they reach training pipelines Track and report data quality metrics across the entire Data Factory Work towards full provenance: training samples should be traceable back to their source geometry and simulation configuration Integration and Delivery Deliver validated datasets to downstream ML training infrastructure in formats optimized for efficient data loading Design data versioning and cataloging systems that support reproducible training runs Work closely with ML Infrastructure Engineers to ensure smooth handoff between data production and model training Support multi-dataset training workflows What you bring to the table Ability to scope and effectively deliver projects, prioritising activity as needed. Problem solving skills and the ability to analyse issues, identify causes, and recommend solutions quickly. Excellent collaboration and communication skills, especially in a research setting. You can translate "the model isn't converging" into infrastructure hypotheses and solutions, and can bridge technical abstractions with implementations. 5+ years of experience in data engineering, HPC engineering, or simulation infrastructure. Strong experience with orchestration systems: SLURM, Kubernetes, Temporal Production data pipeline experience: you've built and operated pipelines that process large volumes of data reliably Proficiency in Python for pipeline development and automation Systems engineering fundamentals: Linux, networking, storage systems, performance debugging Experience with cloud infrastructure; ideally CoreWeave or similar GPU/HPC focused clouds Background in HPC for simulation engineering: experience with CFD, FEA, or similar computational workflows (StarCCM+, OpenFOAM, ANSYS, etc.) Experience with geometry processing: mesh manipulation, CAD formats, PyVista Familiarity with scientific data formats: HDF5, VTK, NetCDF, Zarr Data quality engineering experience: validation frameworks, anomaly detection, data observability Ideally Understanding of CFD fundamentals, enough to interpret solver outputs and validation metrics Experience with 3D geometry pipelines (mesh decimation, field interpolation) Familiarity with ML data loading patterns and how training systems consume data What we offer Equity options - share in our success and growth. 10% employer pension contribution - invest in your future. Free office lunches - great food to fuel your workdays. Flexible working - balance your work and life in a way that works for you. Hybrid setup - enjoy our new Shoreditch office while keeping remote flexibility. Enhanced parental leave - support for life's biggest milestones. Private healthcare - comprehensive coverage Personal development - access learning and training to help you grow. Work from anywhere - extend your remote setup to enjoy the sun or reconnect with loved ones. We value diversity and are committed to equal employment opportunity regardless of sex, race, religion, ethnicity, nationality, disability, age, sexual orientation or gender identity. We strongly encourage individuals from groups traditionally underrepresented in tech to apply. To help make a change, we sponsor bright women from disadvantaged backgrounds through their university degrees in science and mathematics. We collect diversity and inclusion data solely for the purpose of monitoring the effectiveness of our equal opportunities policies and ensuring compliance with UK employment and equality legislation. This information is confidential, used only in aggregate form, and will not influence the outcome of your application.
About the role Anthropic's Infrastructure organization is foundational to our mission of developing AI systems that are reliable, interpretable, and steerable. The systems we build determine how quickly we can train new models, how reliably we can run safety experiments, and how effectively we can scale Claude to millions of users - demonstrating that safe, reliable infrastructure and frontier capabilities can go hand in hand. Node Infra owns the full lifecycle of accelerator capacity at Anthropic. We ingest and provision compute from all major CSPs and our own datacenters, stand up and scale clusters from thousands to hundreds of thousands of hosts, and build the health, diagnostics and repair automation that keep every GPU, TPU and Trainium node in the fleet usable and ready to power Anthropic's frontier AI research. Key responsibilities Own the technical strategy and roadmap for node lifecycle management - ingestion, bring up, health checking, and automated repair Drive cross team initiatives to build and scale AI clusters across multiple clouds and accelerator families Design and operate the systems that detect, isolate, and remediate unhealthy hardware automatically, driving up fleet MTBI and minimizing stranded capacity Define infrastructure architecture, ensuring the hardest problems get solved - whether by you directly or by working through others Work closely with cloud providers and internal research/inference/product teams to shape long term compute, data, and infrastructure strategy Establish and evolve operational excellence practices (incident response, postmortem culture, on call) Support the growth of engineers around you through technical mentorship and coaching Minimum qualifications Deep expertise in distributed systems, reliability, and cloud platforms (e.g., Kubernetes, IaC, AWS/GCP/Azure) Strong proficiency in at least one systems language (e.g., Rust, Go, or Python), IaC proficiency with Terraform Hands on experience with machine learning accelerators (GPUs, TPUs, or Trainium) Track record of leading complex, multi quarter technical initiatives that span multiple teams or systems Ability to build alignment across senior stakeholders and communicate effectively at all levels Bachelor's degree or equivalent in a field relevant to the role Preferred qualifications 8+ years of software engineering experience, including time as a technical lead setting direction for a team Experience managing large scale compute infrastructure at hyperscale (10K+ nodes), including capacity management and efficiency Depth in one or more of: Kubernetes internals (scheduler, autoscaler, kubelet, Karpenter), cluster orchestration systems (Mesos, Borg like), or node provisioning pipelines Low level systems experience: kernel, virtualization, device drivers, firmware, or hardware health/diagnostics daemons Familiarity with high performance networking (EFA, RDMA, InfiniBand) for distributed ML workloads Demonstrated ownership of production reliability for high throughput, latency sensitive systems Contributions to relevant open source projects (Kubernetes, Linux kernel, container runtimes, etc.) Skill in quickly understanding systems design tradeoffs and keeping track of rapidly evolving software systems Compensation Annual salary: £325,000 - £485,000 GBP Benefits Competitive compensation and benefits package Optional equity donation matching Generous vacation and parental leave Flexible working hours Lovely office space in San Francisco
04/07/2026
Full time
About the role Anthropic's Infrastructure organization is foundational to our mission of developing AI systems that are reliable, interpretable, and steerable. The systems we build determine how quickly we can train new models, how reliably we can run safety experiments, and how effectively we can scale Claude to millions of users - demonstrating that safe, reliable infrastructure and frontier capabilities can go hand in hand. Node Infra owns the full lifecycle of accelerator capacity at Anthropic. We ingest and provision compute from all major CSPs and our own datacenters, stand up and scale clusters from thousands to hundreds of thousands of hosts, and build the health, diagnostics and repair automation that keep every GPU, TPU and Trainium node in the fleet usable and ready to power Anthropic's frontier AI research. Key responsibilities Own the technical strategy and roadmap for node lifecycle management - ingestion, bring up, health checking, and automated repair Drive cross team initiatives to build and scale AI clusters across multiple clouds and accelerator families Design and operate the systems that detect, isolate, and remediate unhealthy hardware automatically, driving up fleet MTBI and minimizing stranded capacity Define infrastructure architecture, ensuring the hardest problems get solved - whether by you directly or by working through others Work closely with cloud providers and internal research/inference/product teams to shape long term compute, data, and infrastructure strategy Establish and evolve operational excellence practices (incident response, postmortem culture, on call) Support the growth of engineers around you through technical mentorship and coaching Minimum qualifications Deep expertise in distributed systems, reliability, and cloud platforms (e.g., Kubernetes, IaC, AWS/GCP/Azure) Strong proficiency in at least one systems language (e.g., Rust, Go, or Python), IaC proficiency with Terraform Hands on experience with machine learning accelerators (GPUs, TPUs, or Trainium) Track record of leading complex, multi quarter technical initiatives that span multiple teams or systems Ability to build alignment across senior stakeholders and communicate effectively at all levels Bachelor's degree or equivalent in a field relevant to the role Preferred qualifications 8+ years of software engineering experience, including time as a technical lead setting direction for a team Experience managing large scale compute infrastructure at hyperscale (10K+ nodes), including capacity management and efficiency Depth in one or more of: Kubernetes internals (scheduler, autoscaler, kubelet, Karpenter), cluster orchestration systems (Mesos, Borg like), or node provisioning pipelines Low level systems experience: kernel, virtualization, device drivers, firmware, or hardware health/diagnostics daemons Familiarity with high performance networking (EFA, RDMA, InfiniBand) for distributed ML workloads Demonstrated ownership of production reliability for high throughput, latency sensitive systems Contributions to relevant open source projects (Kubernetes, Linux kernel, container runtimes, etc.) Skill in quickly understanding systems design tradeoffs and keeping track of rapidly evolving software systems Compensation Annual salary: £325,000 - £485,000 GBP Benefits Competitive compensation and benefits package Optional equity donation matching Generous vacation and parental leave Flexible working hours Lovely office space in San Francisco