it job board logo
  • Home
  • Find IT Jobs
  • Register CV
  • Career Advice
  • Contact us
  • Employers
    • Register as Employer
    • Pricing Plans
  • Recruiting? Post a job
  • Sign in
  • Sign up
  • Home
  • Find IT Jobs
  • Register CV
  • Career Advice
  • Contact us
  • Employers
    • Register as Employer
    • Pricing Plans
Sorry, that job is no longer available. Here are some results that may be similar to the job you were looking for.

38 jobs found

Email me jobs like this
Refine Search
Current Search
linux engineer cloud hpc
Technical Futures Ltd
Senior Software Engineer
Technical Futures Ltd Cambridge, Cambridgeshire
Exciting deeptech start-up, an AI Company focused on optimising complex engineering systems, seeks a bright, forward thinking Senior Software / Backend Engineer to join their dedicated team. With strong software engineering skills in modern languages backed by a good academic history, you ll bring curiosity about scientific computing and care for code quality. Applications are welcomed from Senior level up to Staff level Engineers with experience designing systems and the API between them; with experience of modern Python tooling and packaging; knowledge of Cloud computing or HPC job management; identity and authorization flows (such as OAuth2/OIDC). This cutting-edge technology company, focused on optimizing complex engineering systems, seeks a top class Senior Software / Backend Engineer who will confidently work collaboratively with engineers and scientists to have a real say in how the platform layer is designed. You should bring the following skills and experience: Strong academic background. Could be Software Engineering, Computer Science, Physics or Scientific Software related. Strong software engineering skills in modern languages (Python & Rust ideal). Experience designing systems and the APIs between them services that coordinate work, manage state and handle failure. Focus on code quality. Some / most of the following should support the skills above: Experience of Cloud computing or HPC job management. Identity and authorization flows such as OIDC/OAuth2. Deploying containerized services on Linux. Infrastructure as code (such as Ansible, OpenTofu). Modern Python tooling and packaging. Data storage and pipelines, including large simulation results and model serving. API s, REST, WebSockets. Hybrid working available (3 days office /2 WFH), a very generous base salary dependent on your level of skills and experience and benefits to include Shares, 30 days holiday + time off between Xmas and New Year, Private healthcare, Pension Plan, Life Assurance and much more.
22/07/2026
Full time
Exciting deeptech start-up, an AI Company focused on optimising complex engineering systems, seeks a bright, forward thinking Senior Software / Backend Engineer to join their dedicated team. With strong software engineering skills in modern languages backed by a good academic history, you ll bring curiosity about scientific computing and care for code quality. Applications are welcomed from Senior level up to Staff level Engineers with experience designing systems and the API between them; with experience of modern Python tooling and packaging; knowledge of Cloud computing or HPC job management; identity and authorization flows (such as OAuth2/OIDC). This cutting-edge technology company, focused on optimizing complex engineering systems, seeks a top class Senior Software / Backend Engineer who will confidently work collaboratively with engineers and scientists to have a real say in how the platform layer is designed. You should bring the following skills and experience: Strong academic background. Could be Software Engineering, Computer Science, Physics or Scientific Software related. Strong software engineering skills in modern languages (Python & Rust ideal). Experience designing systems and the APIs between them services that coordinate work, manage state and handle failure. Focus on code quality. Some / most of the following should support the skills above: Experience of Cloud computing or HPC job management. Identity and authorization flows such as OIDC/OAuth2. Deploying containerized services on Linux. Infrastructure as code (such as Ansible, OpenTofu). Modern Python tooling and packaging. Data storage and pipelines, including large simulation results and model serving. API s, REST, WebSockets. Hybrid working available (3 days office /2 WFH), a very generous base salary dependent on your level of skills and experience and benefits to include Shares, 30 days holiday + time off between Xmas and New Year, Private healthcare, Pension Plan, Life Assurance and much more.
Technical Futures Ltd
Backend Engineer
Technical Futures Ltd Cambridge, Cambridgeshire
Exciting deeptech start-up seeks a bright, forward thinking Backend Engineer to join their dedicated team. With strong software engineering skills in modern languages (to include Python and Rust) backed by a good academic history, you ll bring curiosity about scientific computing and care for code quality. Applications are welcomed from Mid level up to Senior level Engineers with experience designing systems and the API between them; with experience of modern Python tooling and packaging, some knowledge of Cloud computing or HPC job management (such as Slurm), identity and authorization flows (such as OAuth2/OIDC). This cutting-edge technology company, focused on optimizing complex engineering systems, seeks a top class Backend Engineer who will confidently work collaboratively with engineers and scientists to have a real say in how the platform layer is designed. You should bring the following skills and experience: Strong academic background. Could be Software Engineering, Computer Science, Physics or Scientific Software related. Strong software engineering skills in modern languages (Python & Rust ideal). Experience designing systems and the APIs between them services that coordinate work, manage state and handle failure. Focus on code quality. Some / most of the following should support the skills above: Experience of Cloud computing or HPC job management (such as Slurm). Identity and authorization flows such as OIDC/OAuth2. Deploying containerized services on Linux (such as Podman). Infrastructure as code (such as Ansible, OpenTofu). Modern Python tooling and packaging such as uv. Data storage and pipelines, including large simulation results and model serving. API, REST, WebSockets such as FastAPI. Hybrid working available (3 days office /2 WFH), a very generous base salary dependent on your level of skills and experience and benefits to include Shares, 30 days holiday + time off between Xmas and New Year, Private healthcare, Pension Plan, Life Assurance and much more.
22/07/2026
Full time
Exciting deeptech start-up seeks a bright, forward thinking Backend Engineer to join their dedicated team. With strong software engineering skills in modern languages (to include Python and Rust) backed by a good academic history, you ll bring curiosity about scientific computing and care for code quality. Applications are welcomed from Mid level up to Senior level Engineers with experience designing systems and the API between them; with experience of modern Python tooling and packaging, some knowledge of Cloud computing or HPC job management (such as Slurm), identity and authorization flows (such as OAuth2/OIDC). This cutting-edge technology company, focused on optimizing complex engineering systems, seeks a top class Backend Engineer who will confidently work collaboratively with engineers and scientists to have a real say in how the platform layer is designed. You should bring the following skills and experience: Strong academic background. Could be Software Engineering, Computer Science, Physics or Scientific Software related. Strong software engineering skills in modern languages (Python & Rust ideal). Experience designing systems and the APIs between them services that coordinate work, manage state and handle failure. Focus on code quality. Some / most of the following should support the skills above: Experience of Cloud computing or HPC job management (such as Slurm). Identity and authorization flows such as OIDC/OAuth2. Deploying containerized services on Linux (such as Podman). Infrastructure as code (such as Ansible, OpenTofu). Modern Python tooling and packaging such as uv. Data storage and pipelines, including large simulation results and model serving. API, REST, WebSockets such as FastAPI. Hybrid working available (3 days office /2 WFH), a very generous base salary dependent on your level of skills and experience and benefits to include Shares, 30 days holiday + time off between Xmas and New Year, Private healthcare, Pension Plan, Life Assurance and much more.
PassFort
Staff/Senior Software DevOps Engineer
PassFort
About OLIX AI is growing faster than any technology in history and the explosion in demand has created a massive infrastructure gap; we can no longer build chips or power stations fast enough to keep up. The industry is still leaning on a ten-year-old hardware blueprint that has reached its limit. A new paradigm that is faster and more efficient will be the biggest economic opportunity of the next century and create the most important company of the next decade. The OLIX Decode Accelerator 1 (DX-1) is the first accelerator architected specifically for decode. Rack scale co design of logic, data movement, packaging, optics and interconnect enables a step change in system level performance. The Role We're searching for a Staff/Senior Software DevOps Engineer to own the build, test, and CI flows that the entire DX-1 software stack - compiler, runtime, simulator, and framework integration - depends on. Ours is a large test suite that asserts token exact correctness against golden references, and it has to run across scarce, expensive resources that span both simulation compute and hardware in the loop testing-simulator and emulator boxes alongside DX 1 and prototype platform boards. Your mission is to keep that system fast, trustworthy, observable, and affordable as the test suite, the team, and the resource pool all grow. This is a build and test role, not product serving SRE. You'll work where CI, the runner fleet, and the test hardware meet, partnering closely with the infrastructure, compiler, runtime, simulator, and modelling teams. At Senior/Staff level your impact is the velocity of every engineer who depends on this system - how fast they get a trustworthy signal, how rarely they wait on a machine or a flaky run, and how much they can self serve without coming to you. That leverage - the standards, platforms, and shared resource model others build on - is what we're hiring for, far more than any single system you ship. Responsibilities Own the Build & Test Pipelines: Design, build, and own CI pipelines and test execution - PR, merge, and nightly lanes - that gates the entire software stack, balancing fast feedback against coverage and cost. Scale Test Execution: Split a large, slow suite into staged lanes, parallelise it with real test isolation, and cache aggressively with content addressed keys, so feedback stays fast and cheap as the suite and the team grow rather than getting solved by throwing machines at it. Manage the Fleet & Scarce Resources: Run CI across a heterogeneous fleet of cloud and self hosted machines, and give the team fair, monitored, fail fast shared access to scarce and expensive hardware - keeping it reliable, well utilised, and never a silent bottleneck. Build the Performance & Readiness Signal: Stand up perf regression baselines the team trusts (pinned hardware, rolling baselines, sound metric aggregation, determinism), and turn CI and test signal into CI health and product readiness dashboards that drive real decisions. Own Software Observability: Choose the metrics store that scales to many series on daily runs with long lived history, making dashboards for observable software. Set Standards: Define the flows that keep builds and test runs hermetic and reproducible; and make the system fail closed and contain the blast radius when something is misconfigured or a job is untrusted. Skills & Experience Experience build/test infrastructure, CI/CD, developer productivity, or large scale systems / release engineering, with demonstrated end to end ownership of a big test or CI system. Scaling a large test suite: staged lanes, parallelism with real isolation, content addressed caching, and keeping feedback fast and affordable as the suite grows. Fleet and scarce resource management: heterogeneous CI runner fleets across cloud and on prem (VMs, containers, and bare metal hardware), and giving a team shared, monitored access to scarce or expensive resources - custom accelerators, FPGA/prototype rigs, lab hardware, or contended compute - including reservations, remote access, hardware in the loop testing, and artifact portability across heterogeneous hosts. Performance analysis support: trustworthy regression baselines, determinism and noise handling, sound metric aggregation (geometric vs arithmetic mean vs median), and fast attribution and bisection. Observability and metrics platforms: CI health and readiness dashboards, a metrics store that scales to long lived, high cardinality time series, and self serve access for engineers. Strong scripting and systems programming (e.g. Python plus a systems language), and fluency with containers, Linux, and cloud infrastructure (AWS or similar). A reproducible, safe by default mindset: hermetic builds, least privilege, fail closed defaults, and blast radius control. Excellent communication and the ability to align and influence cross functional teams (compiler, runtime, modelling) without relying on formal authority. Bachelor's degree or higher in computer science, electrical engineering, mathematics, or a related field. Nice to Have GitHub Actions or comparable CI at scale; scaling CI runner fleets on cloud infrastructure (e.g. AWS); hardware in the loop or lab automation for custom silicon or FPGA bring up; time series and observability stacks (Prometheus/Grafana, Datadog, columnar warehouses). Adjacent depth is welcome: HPC / cluster batch scheduling, release engineering, or developer productivity platforms. Compensation & Equity Competitive Salary: Commensurate with your experience, skills, and location Equity & Ownership: Meaningful stock options. You're not just joining the mission; you're owning a piece of it Proximity Bonus: We value your time. To minimise your commute and maximise your life, we offer an annual Living Local Bonus if your residence is within 20 minutes of the office Retirement Benefits: Employer contributed retirement plans to help you build long term financial security. Due to U.S. export control regulations, candidates' eligibility to work at OLIX depends on their most recent citizenship or permanent residency status. We are generally unable to consider applicants whose most recent citizenship or permanent residence is in certain restricted countries (currently including Iran, North Korea, Syria, Cuba, Russia, Belarus, China, Hong Kong, Macau, and Venezuela). Applicants who have subsequently obtained citizenship or permanent residency in another country not subject to these restrictions may still be eligible.
22/07/2026
Full time
About OLIX AI is growing faster than any technology in history and the explosion in demand has created a massive infrastructure gap; we can no longer build chips or power stations fast enough to keep up. The industry is still leaning on a ten-year-old hardware blueprint that has reached its limit. A new paradigm that is faster and more efficient will be the biggest economic opportunity of the next century and create the most important company of the next decade. The OLIX Decode Accelerator 1 (DX-1) is the first accelerator architected specifically for decode. Rack scale co design of logic, data movement, packaging, optics and interconnect enables a step change in system level performance. The Role We're searching for a Staff/Senior Software DevOps Engineer to own the build, test, and CI flows that the entire DX-1 software stack - compiler, runtime, simulator, and framework integration - depends on. Ours is a large test suite that asserts token exact correctness against golden references, and it has to run across scarce, expensive resources that span both simulation compute and hardware in the loop testing-simulator and emulator boxes alongside DX 1 and prototype platform boards. Your mission is to keep that system fast, trustworthy, observable, and affordable as the test suite, the team, and the resource pool all grow. This is a build and test role, not product serving SRE. You'll work where CI, the runner fleet, and the test hardware meet, partnering closely with the infrastructure, compiler, runtime, simulator, and modelling teams. At Senior/Staff level your impact is the velocity of every engineer who depends on this system - how fast they get a trustworthy signal, how rarely they wait on a machine or a flaky run, and how much they can self serve without coming to you. That leverage - the standards, platforms, and shared resource model others build on - is what we're hiring for, far more than any single system you ship. Responsibilities Own the Build & Test Pipelines: Design, build, and own CI pipelines and test execution - PR, merge, and nightly lanes - that gates the entire software stack, balancing fast feedback against coverage and cost. Scale Test Execution: Split a large, slow suite into staged lanes, parallelise it with real test isolation, and cache aggressively with content addressed keys, so feedback stays fast and cheap as the suite and the team grow rather than getting solved by throwing machines at it. Manage the Fleet & Scarce Resources: Run CI across a heterogeneous fleet of cloud and self hosted machines, and give the team fair, monitored, fail fast shared access to scarce and expensive hardware - keeping it reliable, well utilised, and never a silent bottleneck. Build the Performance & Readiness Signal: Stand up perf regression baselines the team trusts (pinned hardware, rolling baselines, sound metric aggregation, determinism), and turn CI and test signal into CI health and product readiness dashboards that drive real decisions. Own Software Observability: Choose the metrics store that scales to many series on daily runs with long lived history, making dashboards for observable software. Set Standards: Define the flows that keep builds and test runs hermetic and reproducible; and make the system fail closed and contain the blast radius when something is misconfigured or a job is untrusted. Skills & Experience Experience build/test infrastructure, CI/CD, developer productivity, or large scale systems / release engineering, with demonstrated end to end ownership of a big test or CI system. Scaling a large test suite: staged lanes, parallelism with real isolation, content addressed caching, and keeping feedback fast and affordable as the suite grows. Fleet and scarce resource management: heterogeneous CI runner fleets across cloud and on prem (VMs, containers, and bare metal hardware), and giving a team shared, monitored access to scarce or expensive resources - custom accelerators, FPGA/prototype rigs, lab hardware, or contended compute - including reservations, remote access, hardware in the loop testing, and artifact portability across heterogeneous hosts. Performance analysis support: trustworthy regression baselines, determinism and noise handling, sound metric aggregation (geometric vs arithmetic mean vs median), and fast attribution and bisection. Observability and metrics platforms: CI health and readiness dashboards, a metrics store that scales to long lived, high cardinality time series, and self serve access for engineers. Strong scripting and systems programming (e.g. Python plus a systems language), and fluency with containers, Linux, and cloud infrastructure (AWS or similar). A reproducible, safe by default mindset: hermetic builds, least privilege, fail closed defaults, and blast radius control. Excellent communication and the ability to align and influence cross functional teams (compiler, runtime, modelling) without relying on formal authority. Bachelor's degree or higher in computer science, electrical engineering, mathematics, or a related field. Nice to Have GitHub Actions or comparable CI at scale; scaling CI runner fleets on cloud infrastructure (e.g. AWS); hardware in the loop or lab automation for custom silicon or FPGA bring up; time series and observability stacks (Prometheus/Grafana, Datadog, columnar warehouses). Adjacent depth is welcome: HPC / cluster batch scheduling, release engineering, or developer productivity platforms. Compensation & Equity Competitive Salary: Commensurate with your experience, skills, and location Equity & Ownership: Meaningful stock options. You're not just joining the mission; you're owning a piece of it Proximity Bonus: We value your time. To minimise your commute and maximise your life, we offer an annual Living Local Bonus if your residence is within 20 minutes of the office Retirement Benefits: Employer contributed retirement plans to help you build long term financial security. Due to U.S. export control regulations, candidates' eligibility to work at OLIX depends on their most recent citizenship or permanent residency status. We are generally unable to consider applicants whose most recent citizenship or permanent residence is in certain restricted countries (currently including Iran, North Korea, Syria, Cuba, Russia, Belarus, China, Hong Kong, Macau, and Venezuela). Applicants who have subsequently obtained citizenship or permanent residency in another country not subject to these restrictions may still be eligible.
Principal Private Cloud Engineer
Arm Limited Cambridge, Cambridgeshire
Job Description: You will define and lead the technical strategy for a private cloud platform based on OpenStack, delivering scalable and reliable infrastructure services for engineering teams. The platform underpins large-scale engineering workloads and is central to the organisation's infrastructure strategy. The platform is to be built using modern platform engineering principles, with strong emphasis on automation, APIs, and developer experience. You will operate across teams and functions to drive architectural direction and platform evolution-the remit of the team is to develop the capabilities of the platform as an active participant within the OpenStack community rather than simply consuming and running upstream projects. Responsibilities: Define and design the long term architecture and strategy for multi tenant Private Cloud. Lead the integration of OpenStack across compute, networking and storage systems. Develop standards for Infrastructure as Code (Terraform), APIs and automation. Drive reliability, scalability and performance improvements across all platform components. Provide technical leadership in collaboration with multiple teams and domains. Guide platform engineering practices including CI/CD, telemetry and observability. Solve complex cross system challenges in distributed environments and HPC. Champion and enable seamless adoption of the new platform. Partner with the Internal Developer Platform teams to ensure a first class self service experience. Identify and improve workflow friction across multi disciplinary teams. Collaborate with the OSS community to drive upstream feature adoption. Required Skills and Experience: Deep expertise in the OpenStack ecosystem, Linux and distributed systems. Strong software engineering background (Python, APIs, systems design). Extensive experience with Ansible, Terraform and infrastructure automation. Experience with large scale or multi region systems and operating at significant scale. Deep understanding of networking, storage and Linux systems both physical and software defined. Proven track record in influencing technical direction across teams. Nice To Have Skills and Experience: Experience with Kubernetes and cloud native platforms. Low level debugging, performance and tuning engineering. Full stack expertise. Exposure to IDPs such as backstage, Roadie, etc. In Return: You will shape the direction of a critical engineering platform. You will influence organisation wide technical decisions. We provide an environment focused on impact, innovation and teamwork. Accommodations at Arm: If you need accommodation during the recruitment process, email . All requests will be handled confidentially. Equal Opportunities at Arm: Arm is an equal opportunity employer committed to a diverse and inclusive workplace. We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, disability or veteran status. Hybrid Working at Arm: Arm's hybrid approach centers on flexibility. Your hybrid working pattern will be determined by you and your team; details will be shared upon application.
21/07/2026
Full time
Job Description: You will define and lead the technical strategy for a private cloud platform based on OpenStack, delivering scalable and reliable infrastructure services for engineering teams. The platform underpins large-scale engineering workloads and is central to the organisation's infrastructure strategy. The platform is to be built using modern platform engineering principles, with strong emphasis on automation, APIs, and developer experience. You will operate across teams and functions to drive architectural direction and platform evolution-the remit of the team is to develop the capabilities of the platform as an active participant within the OpenStack community rather than simply consuming and running upstream projects. Responsibilities: Define and design the long term architecture and strategy for multi tenant Private Cloud. Lead the integration of OpenStack across compute, networking and storage systems. Develop standards for Infrastructure as Code (Terraform), APIs and automation. Drive reliability, scalability and performance improvements across all platform components. Provide technical leadership in collaboration with multiple teams and domains. Guide platform engineering practices including CI/CD, telemetry and observability. Solve complex cross system challenges in distributed environments and HPC. Champion and enable seamless adoption of the new platform. Partner with the Internal Developer Platform teams to ensure a first class self service experience. Identify and improve workflow friction across multi disciplinary teams. Collaborate with the OSS community to drive upstream feature adoption. Required Skills and Experience: Deep expertise in the OpenStack ecosystem, Linux and distributed systems. Strong software engineering background (Python, APIs, systems design). Extensive experience with Ansible, Terraform and infrastructure automation. Experience with large scale or multi region systems and operating at significant scale. Deep understanding of networking, storage and Linux systems both physical and software defined. Proven track record in influencing technical direction across teams. Nice To Have Skills and Experience: Experience with Kubernetes and cloud native platforms. Low level debugging, performance and tuning engineering. Full stack expertise. Exposure to IDPs such as backstage, Roadie, etc. In Return: You will shape the direction of a critical engineering platform. You will influence organisation wide technical decisions. We provide an environment focused on impact, innovation and teamwork. Accommodations at Arm: If you need accommodation during the recruitment process, email . All requests will be handled confidentially. Equal Opportunities at Arm: Arm is an equal opportunity employer committed to a diverse and inclusive workplace. We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, disability or veteran status. Hybrid Working at Arm: Arm's hybrid approach centers on flexibility. Your hybrid working pattern will be determined by you and your team; details will be shared upon application.
HPC Services Team Leader
Viridiengroup Haywards Heath, Sussex
HPC Services Team LeaderApplyremote type: On-sitelocations: Haywards Heath, United Kingdomtime type: Full timeposted on: Posted 2 Days Agojob requisition id: JR101294Viridien () is an advanced technology, digital and Earth data company that pushes the boundaries of science for a more prosperous and sustainable future. With our ingenuity, drive and deep curiosity we discover new insights, innovations, and solutions that efficiently and responsibly resolve complex natural resource, digital, energy transition and infrastructure challenges. Job Summary We are looking for a HPC Services Team Leader to join our global HPC team!Reporting to the Data Center Manager, you will take on a key role in our organization and make a significant impact on our HPC and Cloud environment. In this role, you will work closely with our team to develop our DevOps environment and lead the transition from older IT Ops approaches into a more Agile way of working.You will be responsible for demonstrating best practices and coaching and developing our existing staff. As a senior member of our team, you will play a crucial role in shaping our technology infrastructure and ensuring its smooth operation. With a proven track record across a wide range of activities connected to HPC and Cloud, you will be at the forefront of driving innovation and keeping our organization at the cutting edge of technology.As a self-starter with a deep technical background, you will spearhead the upkeep and evolution of our systems. We need a leader who can credibly navigate the space between technical departments and business stakeholders to drive our infrastructure forward. Key Responsibilities Mentoring and coaching team members. Team building. Challenge, set goals and motivate self-improvement within team. Decision making. Empower team members with skills to improve confidence and technical knowledge. Create engaging and pleasant work environment. Provide support to the Data Centre manager for day-to-day operations. Improve and develop our systems, technology, and infrastructure, alongside providing third line technical support; Make sure the operational maintenance model, and the tools used, are efficient and well-designed. Understanding the client's needs and converting this into technical solutions is important as well as the continued stability, availability, and performance of the platforms. Qualification Degree in any of the following disciplines: Computer Science, Computer Engineering, Computer Information Systems or related computing subjects. Key Skills & Experience Essential 3 - 5 years of leadership experience. Minimum of five years' experience in relevant fields. Extensive levels of Linux administration, preferably in an HPC environment Good experience with Agile Project Management Knowledge in FAI, Puppet, Ansible and Zabbix ITIL Foundation level certification Fast and effective problem-solving skills and a methodical approach to work An enthusiastic attitude towards learning and flexibility to adapt to new challenges or changes in direction The ability to define and manage project deadlines Effective communication Desirable High level DevOps methodologies understanding Knowledge of some of the following: Ansible, OpenStack, Kubernetes, CI/CD, Docker, etc. Scripting and automation Cloud administration experience Virtualization knowledge and experience Experience in hardware maintenance (Storage/CPU/GPU) Why work with us? Competitive salary commensurate with experience Highly attractive bonus scheme Initial 22 days annual leave with future increases, complemented by a flexible buying and selling holiday program Company pension with generous employer contribution Wellbeing Unmind app - puts you in control of your mental health A flexible benefits platform with numerous discount schemes - gym membership, restaurants, cinema tickets, and much more! Regular social club events, spontaneous reward events throughout the year Cycle purchase scheme Flexible Private Medical & Dental care programmes Sponsorship of visas/comprehensive relocation packages Bank Holiday Swap - our holiday swap program allows you to change it for another day of your choice! Relaxed dress code policy Onsite Gym Facilities Learning and Development At Viridien, we foster a culture of continuous learning and provide tailored training programs through our Learning Hub, designed to enhance technical, commercial, and personal growth. We Care about the Environment We encourage and actively support a strong sense of community, through volunteering and various company initiatives, as well as a strong company commitment to protecting our environment through sustainable solutions, energy saving and waste reduction enterprises. Our Hiring Process At Viridien, we are committed to delivering a respectful, inclusive, and transparent recruitment experience.Due to the high volume of applications we receive, we may not be able to provide individual feedback to every applicant. Only candidates whose qualifications closely match the role criteria will be contacted for an interview. We do, however, aim to share personalized feedback with those who progress to the first round of interviews and beyond.We are also dedicated to ensuring that our hiring process accessible to all. If you require any reasonable adjustments to fully participate in the application or interview stages, please don't hesitate to contact your recruiter directly.We see things differently. Diversity fuels our innovation, we value the unique ways in which we differ, and we are committed to equal employment opportunities for all professionals.
18/07/2026
Full time
HPC Services Team LeaderApplyremote type: On-sitelocations: Haywards Heath, United Kingdomtime type: Full timeposted on: Posted 2 Days Agojob requisition id: JR101294Viridien () is an advanced technology, digital and Earth data company that pushes the boundaries of science for a more prosperous and sustainable future. With our ingenuity, drive and deep curiosity we discover new insights, innovations, and solutions that efficiently and responsibly resolve complex natural resource, digital, energy transition and infrastructure challenges. Job Summary We are looking for a HPC Services Team Leader to join our global HPC team!Reporting to the Data Center Manager, you will take on a key role in our organization and make a significant impact on our HPC and Cloud environment. In this role, you will work closely with our team to develop our DevOps environment and lead the transition from older IT Ops approaches into a more Agile way of working.You will be responsible for demonstrating best practices and coaching and developing our existing staff. As a senior member of our team, you will play a crucial role in shaping our technology infrastructure and ensuring its smooth operation. With a proven track record across a wide range of activities connected to HPC and Cloud, you will be at the forefront of driving innovation and keeping our organization at the cutting edge of technology.As a self-starter with a deep technical background, you will spearhead the upkeep and evolution of our systems. We need a leader who can credibly navigate the space between technical departments and business stakeholders to drive our infrastructure forward. Key Responsibilities Mentoring and coaching team members. Team building. Challenge, set goals and motivate self-improvement within team. Decision making. Empower team members with skills to improve confidence and technical knowledge. Create engaging and pleasant work environment. Provide support to the Data Centre manager for day-to-day operations. Improve and develop our systems, technology, and infrastructure, alongside providing third line technical support; Make sure the operational maintenance model, and the tools used, are efficient and well-designed. Understanding the client's needs and converting this into technical solutions is important as well as the continued stability, availability, and performance of the platforms. Qualification Degree in any of the following disciplines: Computer Science, Computer Engineering, Computer Information Systems or related computing subjects. Key Skills & Experience Essential 3 - 5 years of leadership experience. Minimum of five years' experience in relevant fields. Extensive levels of Linux administration, preferably in an HPC environment Good experience with Agile Project Management Knowledge in FAI, Puppet, Ansible and Zabbix ITIL Foundation level certification Fast and effective problem-solving skills and a methodical approach to work An enthusiastic attitude towards learning and flexibility to adapt to new challenges or changes in direction The ability to define and manage project deadlines Effective communication Desirable High level DevOps methodologies understanding Knowledge of some of the following: Ansible, OpenStack, Kubernetes, CI/CD, Docker, etc. Scripting and automation Cloud administration experience Virtualization knowledge and experience Experience in hardware maintenance (Storage/CPU/GPU) Why work with us? Competitive salary commensurate with experience Highly attractive bonus scheme Initial 22 days annual leave with future increases, complemented by a flexible buying and selling holiday program Company pension with generous employer contribution Wellbeing Unmind app - puts you in control of your mental health A flexible benefits platform with numerous discount schemes - gym membership, restaurants, cinema tickets, and much more! Regular social club events, spontaneous reward events throughout the year Cycle purchase scheme Flexible Private Medical & Dental care programmes Sponsorship of visas/comprehensive relocation packages Bank Holiday Swap - our holiday swap program allows you to change it for another day of your choice! Relaxed dress code policy Onsite Gym Facilities Learning and Development At Viridien, we foster a culture of continuous learning and provide tailored training programs through our Learning Hub, designed to enhance technical, commercial, and personal growth. We Care about the Environment We encourage and actively support a strong sense of community, through volunteering and various company initiatives, as well as a strong company commitment to protecting our environment through sustainable solutions, energy saving and waste reduction enterprises. Our Hiring Process At Viridien, we are committed to delivering a respectful, inclusive, and transparent recruitment experience.Due to the high volume of applications we receive, we may not be able to provide individual feedback to every applicant. Only candidates whose qualifications closely match the role criteria will be contacted for an interview. We do, however, aim to share personalized feedback with those who progress to the first round of interviews and beyond.We are also dedicated to ensuring that our hiring process accessible to all. If you require any reasonable adjustments to fully participate in the application or interview stages, please don't hesitate to contact your recruiter directly.We see things differently. Diversity fuels our innovation, we value the unique ways in which we differ, and we are committed to equal employment opportunities for all professionals.
Software Infrastructure Engineer
Cerebras Bristol, Gloucestershire
About Graphcore Graphcore is one of the world's leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world's most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore's teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Summary Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed. The Team The Software Infrastructure team provides critical platforms and services for software development teams across the business. Our responsibilities include managing the CI platform and services, build engineering, component integration, and packaging and release systems. We operate in squads, fostering a culture of service ownership and empowerment for our engineers. We focus on long term engineering solutions and strive to eliminate toil wherever possible. Responsibilities and Duties Develop, own, and maintain tools and services to support the software build and release process Deploy and maintain services with Kubernetes and Docker Manage our Cloud Infrastructure using tools such as Terraform Candidate Profile Essential Knowledge of Python/Go/C++ (or similar language) Experience deploying services in the cloud (AWS preferred) Deep understanding of Linux environments Native user of CI/CD for production deployments Experience with Infrastructure as Code (IaC) tools (e.g. Terraform/OpenTofu) Desirable Experience using Kubernetes (k8s) or OpenStack Experience with GitHub Actions Experience with build tools (e.g. CMake) Experience with modern observability tooling (e.g. Prometheus) Experience with Grafana Benefits In addition to a competitive salary, Graphcore offers flexible working, a generous annual leave policy, private medical insurance and health cash plan, a dental plan, pension (matched up to 5%), life assurance and income protection. We have a generous parental leave policy and an employee assistance programme (which includes health, mental wellbeing, and bereavement support). We offer a range of healthy food and snacks at our central Bristol office and have our own barista bar! We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us.
18/07/2026
Full time
About Graphcore Graphcore is one of the world's leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world's most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore's teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Summary Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed. The Team The Software Infrastructure team provides critical platforms and services for software development teams across the business. Our responsibilities include managing the CI platform and services, build engineering, component integration, and packaging and release systems. We operate in squads, fostering a culture of service ownership and empowerment for our engineers. We focus on long term engineering solutions and strive to eliminate toil wherever possible. Responsibilities and Duties Develop, own, and maintain tools and services to support the software build and release process Deploy and maintain services with Kubernetes and Docker Manage our Cloud Infrastructure using tools such as Terraform Candidate Profile Essential Knowledge of Python/Go/C++ (or similar language) Experience deploying services in the cloud (AWS preferred) Deep understanding of Linux environments Native user of CI/CD for production deployments Experience with Infrastructure as Code (IaC) tools (e.g. Terraform/OpenTofu) Desirable Experience using Kubernetes (k8s) or OpenStack Experience with GitHub Actions Experience with build tools (e.g. CMake) Experience with modern observability tooling (e.g. Prometheus) Experience with Grafana Benefits In addition to a competitive salary, Graphcore offers flexible working, a generous annual leave policy, private medical insurance and health cash plan, a dental plan, pension (matched up to 5%), life assurance and income protection. We have a generous parental leave policy and an employee assistance programme (which includes health, mental wellbeing, and bereavement support). We offer a range of healthy food and snacks at our central Bristol office and have our own barista bar! We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us.
Platform Engineer - Developer Platform & DevOps Tooling
Viridien Crawley, Sussex
Viridien ( ) is an advanced technology, digital and Earth data company that pushes the boundaries of science for a more prosperous and sustainable future. With our ingenuity, drive and deep curiosity we discover new insights, innovations, and solutions that efficiently and responsibly resolve complex natural resource, digital, energy transition and infrastructure challenges.Looking for a role where curiosity, engineering discipline, and platform innovation come together in a unique HPC environment?Viridien is seeking a Platform Engineer - Developer Platform & DevOps to help create the tools, services, and automation that enable engineering teams to develop, deploy, and operate software across a large-scale HPC platform.This is a platform engineering role that combines software development, DevOps, and platform operations. We are looking for someone with a developer background who wants to grow their platform engineering skills while writing maintainable code, supporting production services, and improving the platforms that engineering teams rely on day to day.You will work across CI/CD systems, GitLab, Kubernetes services, Puppet-managed environments, observability tooling, and OpenStack-backed infrastructure. OpenStack and AI-related experience are useful but not required.About The TeamYou will join a team focused on platform engineering and developer tooling, supporting scalable software development, reliable software delivery, and large-scale compute environments.The team works closely with software engineering, infrastructure, operations, and research teams to build practical tools and platform capabilities that help developers ship and operate software more effectively.QualificationsRequiredSoftware development experience in one or more modern programming languages; Go, C++, Python, and scripting experience are especially useful.Experience building or contributing to tools, services, APIs, platforms, or distributed systems.Familiarity with CI/CD pipelines, automated testing, release processes, or deployment automation.Practical experience with Linux and containers, with exposure to Kubernetes or similar orchestration platforms.Ability to write maintainable code and willingness to learn how to operate it in real environments.Good debugging and problem-solving skills, with an interest in understanding issues across software and platform layers.Comfortable collaborating across development, platform, and operations teams.DesirableExperience with AI-powered CLI tools, reusable agent/AI skills, agent-based or AI-assisted developer workflows, and LLM-based tooling.Experience with observability tools such as Prometheus, Grafana, ELK, Thanos, or similar.Experience with infrastructure automation, infrastructure-as-code, or deployment tooling such as Terraform or Helm.Experience with Puppet, Ansible, or similar configuration management tooling.Exposure to OpenStack, cloud infrastructure, and large-scale compute environments.Familiarity with authentication and access control systems such as OIDC.Experience supporting high-performance, data-intensive, or research-oriented workloads.Why work with us?£40,000-£46,000 per annum depending on experienceHighly attractive bonus schemeInitial 22 days annual leave with future increases, complemented by a flexible buying and selling holiday programCompany pension with generous employer contributionWellbeing Unmind app - puts you in control of your mental healthA flexible benefits platform with numerous discount schemes - gym membership, restaurants, cinema tickets, and much more!Regular social club events, spontaneous reward events throughout the yearCycle purchase schemeFlexible Private Medical & Dental care programmesSponsorship of visas/comprehensive relocation packagesBank Holiday Swap - our holiday swap program allows you to change it for another day of your choice!Relaxed dress code policyLearning and DevelopmentAt Viridien, we foster a culture of continuous learning and provide tailored training programs through our Learning Hub, designed to enhance technical, commercial, and personal growth.We Care About The EnvironmentWe encourage and actively support a strong sense of community, through volunteering and various company initiatives, as well as a strong company commitment to protecting our environment through sustainable solutions, energy saving and waste reduction enterprises.Our Hiring ProcessAt Viridien, we are committed to delivering a respectful, inclusive, and transparent recruitment experience.Due to the high volume of applications we receive, we may not be able to provide individual feedback to every applicant. Only candidates whose qualifications closely match the role criteria will be contacted for an interview. We do, however, aim to share personalized feedback with those who progress to the first round of interviews and beyond.We are also dedicated to ensuring that our hiring process accessible to all. If you require any reasonable adjustments to fully participate in the application or interview stages, please don't hesitate to contact your recruiter directly.We see things differently. Diversity fuels our innovation, we value the unique ways in which we differ, and we are committed to equal employment opportunities for all professionals.
16/07/2026
Full time
Viridien ( ) is an advanced technology, digital and Earth data company that pushes the boundaries of science for a more prosperous and sustainable future. With our ingenuity, drive and deep curiosity we discover new insights, innovations, and solutions that efficiently and responsibly resolve complex natural resource, digital, energy transition and infrastructure challenges.Looking for a role where curiosity, engineering discipline, and platform innovation come together in a unique HPC environment?Viridien is seeking a Platform Engineer - Developer Platform & DevOps to help create the tools, services, and automation that enable engineering teams to develop, deploy, and operate software across a large-scale HPC platform.This is a platform engineering role that combines software development, DevOps, and platform operations. We are looking for someone with a developer background who wants to grow their platform engineering skills while writing maintainable code, supporting production services, and improving the platforms that engineering teams rely on day to day.You will work across CI/CD systems, GitLab, Kubernetes services, Puppet-managed environments, observability tooling, and OpenStack-backed infrastructure. OpenStack and AI-related experience are useful but not required.About The TeamYou will join a team focused on platform engineering and developer tooling, supporting scalable software development, reliable software delivery, and large-scale compute environments.The team works closely with software engineering, infrastructure, operations, and research teams to build practical tools and platform capabilities that help developers ship and operate software more effectively.QualificationsRequiredSoftware development experience in one or more modern programming languages; Go, C++, Python, and scripting experience are especially useful.Experience building or contributing to tools, services, APIs, platforms, or distributed systems.Familiarity with CI/CD pipelines, automated testing, release processes, or deployment automation.Practical experience with Linux and containers, with exposure to Kubernetes or similar orchestration platforms.Ability to write maintainable code and willingness to learn how to operate it in real environments.Good debugging and problem-solving skills, with an interest in understanding issues across software and platform layers.Comfortable collaborating across development, platform, and operations teams.DesirableExperience with AI-powered CLI tools, reusable agent/AI skills, agent-based or AI-assisted developer workflows, and LLM-based tooling.Experience with observability tools such as Prometheus, Grafana, ELK, Thanos, or similar.Experience with infrastructure automation, infrastructure-as-code, or deployment tooling such as Terraform or Helm.Experience with Puppet, Ansible, or similar configuration management tooling.Exposure to OpenStack, cloud infrastructure, and large-scale compute environments.Familiarity with authentication and access control systems such as OIDC.Experience supporting high-performance, data-intensive, or research-oriented workloads.Why work with us?£40,000-£46,000 per annum depending on experienceHighly attractive bonus schemeInitial 22 days annual leave with future increases, complemented by a flexible buying and selling holiday programCompany pension with generous employer contributionWellbeing Unmind app - puts you in control of your mental healthA flexible benefits platform with numerous discount schemes - gym membership, restaurants, cinema tickets, and much more!Regular social club events, spontaneous reward events throughout the yearCycle purchase schemeFlexible Private Medical & Dental care programmesSponsorship of visas/comprehensive relocation packagesBank Holiday Swap - our holiday swap program allows you to change it for another day of your choice!Relaxed dress code policyLearning and DevelopmentAt Viridien, we foster a culture of continuous learning and provide tailored training programs through our Learning Hub, designed to enhance technical, commercial, and personal growth.We Care About The EnvironmentWe encourage and actively support a strong sense of community, through volunteering and various company initiatives, as well as a strong company commitment to protecting our environment through sustainable solutions, energy saving and waste reduction enterprises.Our Hiring ProcessAt Viridien, we are committed to delivering a respectful, inclusive, and transparent recruitment experience.Due to the high volume of applications we receive, we may not be able to provide individual feedback to every applicant. Only candidates whose qualifications closely match the role criteria will be contacted for an interview. We do, however, aim to share personalized feedback with those who progress to the first round of interviews and beyond.We are also dedicated to ensuring that our hiring process accessible to all. If you require any reasonable adjustments to fully participate in the application or interview stages, please don't hesitate to contact your recruiter directly.We see things differently. Diversity fuels our innovation, we value the unique ways in which we differ, and we are committed to equal employment opportunities for all professionals.
Linuxrecruit
Platform Engineer
Linuxrecruit
Overview Are you ready for the cloud migration revolution? Many companies have already made the switch, migrating their on-premise legacy infrastructure to one of the many cloud service providers, as it requires less volatile hardware and reduces their overall operational costs. This role gives you the opportunity to join the revolution and get your hands dirty with building out IAS using Terraform and AWS. The scale of the projects that this role offers you is currently unrivalled, as you'd be working on the largest implementation of Kubernetes in the UK. You'd be working within the organisation's "centre of excellence" for platform engineering, so you'd be in good company alongside other best-in-class engineers. To add even more strings to your bow, the company offers you access to an unlimited training budget, which has allowed over 90% of the current platform engineering team to become certified with various technologies. The role has opened due to ongoing success worldwide, leading to several new contracts being taken on by the company within the past few months. This role also offers you the flexibility to work remotely, with no strict working hours which means you can fit your job around the rest of your life rather than the other way round. If you have experience in the mentioned tech above, as well as with building out CI/CD pipelines then don't hesitate to reach out. No CVs are required for an initial conversation. A lover of Linux, automation, containers, bare metal environment? Well, this is the right role for you Role: Technical Account Manager (Observability) As data volumes surge and observability costs climb, this is a rare opportunity to join a high-growth SaaS company at the forefront of modern observability, helping companies tackle these challenges with smarter, more scalable solutions. AI, Compilers, HPC, Low-latency systems, Python/C++, if any of those get you excited, then keep reading Whatever role you are looking for, our team will work with you to understand your unique skills,experience, career goals and aspirations. Kick-start your job search by registering with us today. Searching for new talent? Let's go. Get in touch with us today to find out how we can help scale your team.
15/07/2026
Full time
Overview Are you ready for the cloud migration revolution? Many companies have already made the switch, migrating their on-premise legacy infrastructure to one of the many cloud service providers, as it requires less volatile hardware and reduces their overall operational costs. This role gives you the opportunity to join the revolution and get your hands dirty with building out IAS using Terraform and AWS. The scale of the projects that this role offers you is currently unrivalled, as you'd be working on the largest implementation of Kubernetes in the UK. You'd be working within the organisation's "centre of excellence" for platform engineering, so you'd be in good company alongside other best-in-class engineers. To add even more strings to your bow, the company offers you access to an unlimited training budget, which has allowed over 90% of the current platform engineering team to become certified with various technologies. The role has opened due to ongoing success worldwide, leading to several new contracts being taken on by the company within the past few months. This role also offers you the flexibility to work remotely, with no strict working hours which means you can fit your job around the rest of your life rather than the other way round. If you have experience in the mentioned tech above, as well as with building out CI/CD pipelines then don't hesitate to reach out. No CVs are required for an initial conversation. A lover of Linux, automation, containers, bare metal environment? Well, this is the right role for you Role: Technical Account Manager (Observability) As data volumes surge and observability costs climb, this is a rare opportunity to join a high-growth SaaS company at the forefront of modern observability, helping companies tackle these challenges with smarter, more scalable solutions. AI, Compilers, HPC, Low-latency systems, Python/C++, if any of those get you excited, then keep reading Whatever role you are looking for, our team will work with you to understand your unique skills,experience, career goals and aspirations. Kick-start your job search by registering with us today. Searching for new talent? Let's go. Get in touch with us today to find out how we can help scale your team.
Linuxrecruit
Platform Engineer - Cloud, Kubernetes & CI/CD (Remote)
Linuxrecruit
Overview Are you ready for the cloud migration revolution? Many companies have already made the switch, migrating their on-premise legacy infrastructure to one of the many cloud service providers, as it requires less volatile hardware and reduces their overall operational costs. This role gives you the opportunity to join the revolution and get your hands dirty with building out IAS using Terraform and AWS. The scale of the projects that this role offers you is currently unrivalled, as you'd be working on the largest implementation of Kubernetes in the UK. You'd be working within the organisation's "centre of excellence" for platform engineering, so you'd be in good company alongside other best-in-class engineers. To add even more strings to your bow, the company offers you access to an unlimited training budget, which has allowed over 90% of the current platform engineering team to become certified with various technologies. The role has opened due to ongoing success worldwide, leading to several new contracts being taken on by the company within the past few months. This role also offers you the flexibility to work remotely, with no strict working hours which means you can fit your job around the rest of your life rather than the other way round. If you have experience in the mentioned tech above, as well as with building out CI/CD pipelines then don't hesitate to reach out. No CVs are required for an initial conversation. A lover of Linux, automation, containers, bare metal environment? Well, this is the right role for you Role: Technical Account Manager (Observability) As data volumes surge and observability costs climb, this is a rare opportunity to join a high-growth SaaS company at the forefront of modern observability, helping companies tackle these challenges with smarter, more scalable solutions. AI, Compilers, HPC, Low-latency systems, Python/C++, if any of those get you excited, then keep reading Whatever role you are looking for, our team will work with you to understand your unique skills,experience, career goals and aspirations. Kick-start your job search by registering with us today. Searching for new talent? Let's go. Get in touch with us today to find out how we can help scale your team.
15/07/2026
Full time
Overview Are you ready for the cloud migration revolution? Many companies have already made the switch, migrating their on-premise legacy infrastructure to one of the many cloud service providers, as it requires less volatile hardware and reduces their overall operational costs. This role gives you the opportunity to join the revolution and get your hands dirty with building out IAS using Terraform and AWS. The scale of the projects that this role offers you is currently unrivalled, as you'd be working on the largest implementation of Kubernetes in the UK. You'd be working within the organisation's "centre of excellence" for platform engineering, so you'd be in good company alongside other best-in-class engineers. To add even more strings to your bow, the company offers you access to an unlimited training budget, which has allowed over 90% of the current platform engineering team to become certified with various technologies. The role has opened due to ongoing success worldwide, leading to several new contracts being taken on by the company within the past few months. This role also offers you the flexibility to work remotely, with no strict working hours which means you can fit your job around the rest of your life rather than the other way round. If you have experience in the mentioned tech above, as well as with building out CI/CD pipelines then don't hesitate to reach out. No CVs are required for an initial conversation. A lover of Linux, automation, containers, bare metal environment? Well, this is the right role for you Role: Technical Account Manager (Observability) As data volumes surge and observability costs climb, this is a rare opportunity to join a high-growth SaaS company at the forefront of modern observability, helping companies tackle these challenges with smarter, more scalable solutions. AI, Compilers, HPC, Low-latency systems, Python/C++, if any of those get you excited, then keep reading Whatever role you are looking for, our team will work with you to understand your unique skills,experience, career goals and aspirations. Kick-start your job search by registering with us today. Searching for new talent? Let's go. Get in touch with us today to find out how we can help scale your team.
HPC Services Team Leader
CGG Services (UK) Limited Bolney, Sussex
Company Overview Viridien () is an advanced technology, digital and Earth data company that pushes the boundaries of science for a more prosperous and sustainable future. With our ingenuity, drive and deep curiosity we discover new insights, innovations, and solutions that efficiently and responsibly resolve complex natural resource, digital, energy transition and infrastructure challenges. Job Summary We are looking for a HPC Services Team Leader to join our global HPC team! Reporting to the Data Center Manager, you will play a key role in developing our DevOps environment and leading the transition from legacy IT Ops approaches to an Agile way of working. You will demonstrate best practices, coach and develop staff, shape our technology infrastructure, and ensure its smooth operation. Key Responsibilities Mentor and coach team members, build the team, set goals, and motivate self improvement. Make decisions and empower team members with skills to improve confidence and technical knowledge. Create an engaging and pleasant work environment. Support the Data Centre Manager for day to day operations. Improve and develop our systems, technology, and infrastructure, and provide third line technical support. Ensure the operational maintenance model and tools are efficient and well designed. Understand client needs and convert them into technical solutions, while maintaining stability, availability, and performance of the platforms. Qualifications Degree in Computer Science, Computer Engineering, Computer Information Systems or related computing subjects. 3-5 years of leadership experience and a minimum of five years' experience in relevant fields. Extensive experience with Linux administration, preferably in an HPC environment. Experience with Agile Project Management, FAI, Puppet, Ansible and Zabbix. ITIL Foundation level certification. Fast and effective problem solving skills, methodical approach, enthusiastic learning attitude, and flexibility to adapt to new challenges. Ability to define and manage project deadlines and communicate effectively. Desired Skills & Experience High level DevOps methodology understanding. Knowledge of Ansible, OpenStack, Kubernetes, CI/CD, Docker, etc. Cloud administration experience. Virtualization knowledge. Hardware maintenance (storage, CPU, GPU). Benefits Competitive salary commensurate with experience. Highly attractive bonus scheme. 22 days of annual leave, with future increases. Flexible buying and selling holiday programme. Company pension with generous employer contribution. Wellbeing: Unmind app, flexible benefits platform, discount schemes (gym, restaurants, cinema, etc.). Regular social club events and reward events. Cycle purchase scheme. Flexible Private Medical & Dental care programmes. Sponsorship of visas and relocation packages. Bank holiday swap program. Relaxed dress code. Onsite gym facilities. Learning and development opportunities. Equal Employment Opportunity We are committed to equal employment opportunities for all professionals. We value the unique ways in which we differ and believe diversity fuels our innovation.
15/07/2026
Full time
Company Overview Viridien () is an advanced technology, digital and Earth data company that pushes the boundaries of science for a more prosperous and sustainable future. With our ingenuity, drive and deep curiosity we discover new insights, innovations, and solutions that efficiently and responsibly resolve complex natural resource, digital, energy transition and infrastructure challenges. Job Summary We are looking for a HPC Services Team Leader to join our global HPC team! Reporting to the Data Center Manager, you will play a key role in developing our DevOps environment and leading the transition from legacy IT Ops approaches to an Agile way of working. You will demonstrate best practices, coach and develop staff, shape our technology infrastructure, and ensure its smooth operation. Key Responsibilities Mentor and coach team members, build the team, set goals, and motivate self improvement. Make decisions and empower team members with skills to improve confidence and technical knowledge. Create an engaging and pleasant work environment. Support the Data Centre Manager for day to day operations. Improve and develop our systems, technology, and infrastructure, and provide third line technical support. Ensure the operational maintenance model and tools are efficient and well designed. Understand client needs and convert them into technical solutions, while maintaining stability, availability, and performance of the platforms. Qualifications Degree in Computer Science, Computer Engineering, Computer Information Systems or related computing subjects. 3-5 years of leadership experience and a minimum of five years' experience in relevant fields. Extensive experience with Linux administration, preferably in an HPC environment. Experience with Agile Project Management, FAI, Puppet, Ansible and Zabbix. ITIL Foundation level certification. Fast and effective problem solving skills, methodical approach, enthusiastic learning attitude, and flexibility to adapt to new challenges. Ability to define and manage project deadlines and communicate effectively. Desired Skills & Experience High level DevOps methodology understanding. Knowledge of Ansible, OpenStack, Kubernetes, CI/CD, Docker, etc. Cloud administration experience. Virtualization knowledge. Hardware maintenance (storage, CPU, GPU). Benefits Competitive salary commensurate with experience. Highly attractive bonus scheme. 22 days of annual leave, with future increases. Flexible buying and selling holiday programme. Company pension with generous employer contribution. Wellbeing: Unmind app, flexible benefits platform, discount schemes (gym, restaurants, cinema, etc.). Regular social club events and reward events. Cycle purchase scheme. Flexible Private Medical & Dental care programmes. Sponsorship of visas and relocation packages. Bank holiday swap program. Relaxed dress code. Onsite gym facilities. Learning and development opportunities. Equal Employment Opportunity We are committed to equal employment opportunities for all professionals. We value the unique ways in which we differ and believe diversity fuels our innovation.
IT Graduate Recruitment
High Performance Computer Scientist /HPC Developer
IT Graduate Recruitment
We are looking for an exceptional HPC Engineer to design and optimise some of the most demanding computational systems. This role is focused on solving difficult engineering problems at scale, working with advanced computing infrastructure, distributed systems and performance-critical workloads. We are seeking someone with outstanding technical ability who enjoys understanding systems deeply, from hardware and operating systems through to software optimisation and large-scale computation. About the Role You will work on the design, optimisation and operation of high-performance computing environments. You will be responsible for improving computational efficiency, building reliable infrastructure and solving complex challenges involving large-scale workloads. This is a highly technical role suited to someone with exceptional analytical ability and a passion for pushing the limits of computing performance. Key Responsibilities Design and optimise high-performance computing environments. Build and maintain large-scale Linux-based compute infrastructure. Improve performance through profiling, benchmarking and optimisation. Develop automation tools for managing complex computational systems. Work with distributed computing workloads and parallel processing environments. Diagnose and resolve challenging infrastructure and performance issues. Optimise resource utilisation across compute, storage and networking systems. Develop solutions for reliability, scalability and operational efficiency. Collaborate with engineers and researchers on computationally intensive projects. About You Exceptional academic background from a leading UK university. Degree in Computer Science, Mathematics, Physics, Engineering or a related quantitative discipline. Strong programming ability in Python, C++, C, Rust or Fortran. Deep understanding of Linux systems and computer architecture. Strong knowledge of algorithms, performance optimisation and systems engineering. Ability to analyse complex technical problems and develop elegant solutions. Genuine interest in low-level systems, computing performance and large-scale infrastructure. Strong Signals First-class degree or exceptional academic record. Experience with HPC, supercomputing or large-scale distributed systems. Research experience involving computational workloads. Experience with parallel programming (MPI, OpenMP, CUDA). Knowledge of scheduling systems such as Slurm, PBS or LSF. Contributions to technical projects, open source or research communities. Experience working with advanced computing environments. Academic Focus We are particularly interested in exceptional candidates from the UK's strongest technical universities, including those with backgrounds in Computer Science, Mathematics, Physics and Engineering. We value demonstrated problem-solving ability, technical depth and evidence of tackling genuinely difficult computational challenges over years of experience alone. HPC Engineer - Keyword List HPC Engineer, High Performance Computing Engineer, Research Engineer, Systems Engineer, Infrastructure Engineer, Compute Engineer, Software Engineer, Linux Engineer, Distributed Systems Engineer, Supercomputing, High Performance Computing, HPC, Linux, Unix, Systems Programming, Computer Architecture, Operating Systems, Parallel Computing, Distributed Computing, Cluster Computing, Compute Clusters, Cloud HPC, Large Scale Computing, Scientific Computing, Numerical Computing, Computational Science, Computational Engineering, Performance Engineering, Performance Optimisation, Benchmarking, Profiling, Low Level Programming, C++, C, Python, Fortran, Rust, CUDA, GPU Computing, GPU Acceleration, Parallel Programming, MPI, OpenMP, Multithreading, Concurrency, Algorithms, Data Structures, Systems Design, Kernel Development, Networking, Storage Systems, Distributed Storage, Automation, Infrastructure Automation, Shell Scripting, Bash, Slurm, PBS, LSF, Workload Scheduling, Resource Management, Linux Administration, Server Infrastructure, Cloud Infrastructure, AWS HPC, Azure HPC, Data Processing, Machine Learning Infrastructure, AI Infrastructure, Scientific Research,
14/07/2026
Full time
We are looking for an exceptional HPC Engineer to design and optimise some of the most demanding computational systems. This role is focused on solving difficult engineering problems at scale, working with advanced computing infrastructure, distributed systems and performance-critical workloads. We are seeking someone with outstanding technical ability who enjoys understanding systems deeply, from hardware and operating systems through to software optimisation and large-scale computation. About the Role You will work on the design, optimisation and operation of high-performance computing environments. You will be responsible for improving computational efficiency, building reliable infrastructure and solving complex challenges involving large-scale workloads. This is a highly technical role suited to someone with exceptional analytical ability and a passion for pushing the limits of computing performance. Key Responsibilities Design and optimise high-performance computing environments. Build and maintain large-scale Linux-based compute infrastructure. Improve performance through profiling, benchmarking and optimisation. Develop automation tools for managing complex computational systems. Work with distributed computing workloads and parallel processing environments. Diagnose and resolve challenging infrastructure and performance issues. Optimise resource utilisation across compute, storage and networking systems. Develop solutions for reliability, scalability and operational efficiency. Collaborate with engineers and researchers on computationally intensive projects. About You Exceptional academic background from a leading UK university. Degree in Computer Science, Mathematics, Physics, Engineering or a related quantitative discipline. Strong programming ability in Python, C++, C, Rust or Fortran. Deep understanding of Linux systems and computer architecture. Strong knowledge of algorithms, performance optimisation and systems engineering. Ability to analyse complex technical problems and develop elegant solutions. Genuine interest in low-level systems, computing performance and large-scale infrastructure. Strong Signals First-class degree or exceptional academic record. Experience with HPC, supercomputing or large-scale distributed systems. Research experience involving computational workloads. Experience with parallel programming (MPI, OpenMP, CUDA). Knowledge of scheduling systems such as Slurm, PBS or LSF. Contributions to technical projects, open source or research communities. Experience working with advanced computing environments. Academic Focus We are particularly interested in exceptional candidates from the UK's strongest technical universities, including those with backgrounds in Computer Science, Mathematics, Physics and Engineering. We value demonstrated problem-solving ability, technical depth and evidence of tackling genuinely difficult computational challenges over years of experience alone. HPC Engineer - Keyword List HPC Engineer, High Performance Computing Engineer, Research Engineer, Systems Engineer, Infrastructure Engineer, Compute Engineer, Software Engineer, Linux Engineer, Distributed Systems Engineer, Supercomputing, High Performance Computing, HPC, Linux, Unix, Systems Programming, Computer Architecture, Operating Systems, Parallel Computing, Distributed Computing, Cluster Computing, Compute Clusters, Cloud HPC, Large Scale Computing, Scientific Computing, Numerical Computing, Computational Science, Computational Engineering, Performance Engineering, Performance Optimisation, Benchmarking, Profiling, Low Level Programming, C++, C, Python, Fortran, Rust, CUDA, GPU Computing, GPU Acceleration, Parallel Programming, MPI, OpenMP, Multithreading, Concurrency, Algorithms, Data Structures, Systems Design, Kernel Development, Networking, Storage Systems, Distributed Storage, Automation, Infrastructure Automation, Shell Scripting, Bash, Slurm, PBS, LSF, Workload Scheduling, Resource Management, Linux Administration, Server Infrastructure, Cloud Infrastructure, AWS HPC, Azure HPC, Data Processing, Machine Learning Infrastructure, AI Infrastructure, Scientific Research,
DevOps Engineer
Diffractive Labs
What We're Looking For We are seeking a DevOps Engineer to build and own the infrastructure that underpins our AI driven materials discovery platform. You'll work directly with world renowned ML researchers and software engineers to accelerate real scientific breakthroughs by making model training, experimentation, and deployment fast, reliable, and reproducible. This is a foundational hire. You'll set the patterns others build on. You will be joining a small, highly ambitious team of world renowned engineers, AI researchers, and materials scientists. We move fast and value people who are energised by that. What You'll Do Design, provision, and manage cloud infrastructure (AWS/GCP) using infrastructure as code; Terraform, Pulumi, or equivalent. Own GPU compute environments for model training and inference, including cluster configuration, job scheduling, and cost optimisation. Build and maintain CI/CD pipelines that support rapid model iteration, automated testing, and safe deployments. Support ML workflow orchestration; experiment tracking, training run management, and data pipeline reliability. Ensure reproducibility across research and production environments through containerisation and rigorous environment management. Define monitoring, alerting, and incident response processes so the team can move fast without things silently breaking. Implement security best practices: secrets management, IAM, network segmentation, vulnerability scanning. Build internal tooling and documentation that lets researchers self serve infrastructure without waiting on you. Skills & Qualifications 4+ years in a DevOps, Platform Engineering, or SRE role. Strong proficiency with at least one major cloud provider and its core services (compute, storage, networking, IAM). Hands on experience with infrastructure as code and container orchestration (Kubernetes or equivalent). Solid CI/CD pipeline experience, GitHub Actions, GitLab CI, or similar. Proficient in Python and Bash; comfortable reading and writing code across a polyglot stack. Deep Linux systems knowledge and strong networking fundamentals. A bias for building things properly the first time, even under early stage constraints. Nice to Have Experience with GPU cluster management and ML training workloads (NVIDIA, CUDA, distributed training). Familiarity with MLOps tooling: Experiment tracking (MLflow, Weights & Biases). Workflow orchestration (Airflow, Prefect, Argo). Data versioning (DVC). Background in scientific computing or HPC environments. Prior experience at a deep tech or computational science company. Why Join Us Work directly on infrastructure that enables AI to make real scientific discoveries. Shape how we build from day one, no legacy systems, no inherited mess. Collaborate with world class researchers across materials science and machine learning. Diffractive is building the AI Material Scientist that autonomously learns from real world experimentation to push the boundaries of scientific discovery. We're early, moving fast, and working on problems that genuinely matter. We are a London based company with a flexible approach to how and where you work. We offer competitive salary, generous equity, and benefits. You'll have a real stake in what you build and in the company's overall success. Equal Opportunity Diffractive is an equal opportunities employer. We are committed to creating an inclusive environment for all employees and welcome applications from people of all backgrounds, experiences, and identities. If you require any adjustments or accommodations at any point during the interview process please let us know - we will be happy to help.
11/07/2026
Full time
What We're Looking For We are seeking a DevOps Engineer to build and own the infrastructure that underpins our AI driven materials discovery platform. You'll work directly with world renowned ML researchers and software engineers to accelerate real scientific breakthroughs by making model training, experimentation, and deployment fast, reliable, and reproducible. This is a foundational hire. You'll set the patterns others build on. You will be joining a small, highly ambitious team of world renowned engineers, AI researchers, and materials scientists. We move fast and value people who are energised by that. What You'll Do Design, provision, and manage cloud infrastructure (AWS/GCP) using infrastructure as code; Terraform, Pulumi, or equivalent. Own GPU compute environments for model training and inference, including cluster configuration, job scheduling, and cost optimisation. Build and maintain CI/CD pipelines that support rapid model iteration, automated testing, and safe deployments. Support ML workflow orchestration; experiment tracking, training run management, and data pipeline reliability. Ensure reproducibility across research and production environments through containerisation and rigorous environment management. Define monitoring, alerting, and incident response processes so the team can move fast without things silently breaking. Implement security best practices: secrets management, IAM, network segmentation, vulnerability scanning. Build internal tooling and documentation that lets researchers self serve infrastructure without waiting on you. Skills & Qualifications 4+ years in a DevOps, Platform Engineering, or SRE role. Strong proficiency with at least one major cloud provider and its core services (compute, storage, networking, IAM). Hands on experience with infrastructure as code and container orchestration (Kubernetes or equivalent). Solid CI/CD pipeline experience, GitHub Actions, GitLab CI, or similar. Proficient in Python and Bash; comfortable reading and writing code across a polyglot stack. Deep Linux systems knowledge and strong networking fundamentals. A bias for building things properly the first time, even under early stage constraints. Nice to Have Experience with GPU cluster management and ML training workloads (NVIDIA, CUDA, distributed training). Familiarity with MLOps tooling: Experiment tracking (MLflow, Weights & Biases). Workflow orchestration (Airflow, Prefect, Argo). Data versioning (DVC). Background in scientific computing or HPC environments. Prior experience at a deep tech or computational science company. Why Join Us Work directly on infrastructure that enables AI to make real scientific discoveries. Shape how we build from day one, no legacy systems, no inherited mess. Collaborate with world class researchers across materials science and machine learning. Diffractive is building the AI Material Scientist that autonomously learns from real world experimentation to push the boundaries of scientific discovery. We're early, moving fast, and working on problems that genuinely matter. We are a London based company with a flexible approach to how and where you work. We offer competitive salary, generous equity, and benefits. You'll have a real stake in what you build and in the company's overall success. Equal Opportunity Diffractive is an equal opportunities employer. We are committed to creating an inclusive environment for all employees and welcome applications from people of all backgrounds, experiences, and identities. If you require any adjustments or accommodations at any point during the interview process please let us know - we will be happy to help.
Senior Cloud Network Engineer: Private Cloud & HPC Networking
Cerebras Bristol, Gloucestershire
Cerebras is seeking a Senior Network Engineer in Bristol to develop and deploy high-performance cloud services. This role involves close collaboration with cross-functional teams and hands-on technical work surrounding cloud infrastructure, automation, and performance optimization. Candidates should have significant experience with high-end Ethernet switches, cloud environment management, and strong Linux skills. The position comes with a competitive salary and comprehensive benefits including flexible working and private medical insurance.
11/07/2026
Full time
Cerebras is seeking a Senior Network Engineer in Bristol to develop and deploy high-performance cloud services. This role involves close collaboration with cross-functional teams and hands-on technical work surrounding cloud infrastructure, automation, and performance optimization. Candidates should have significant experience with high-end Ethernet switches, cloud environment management, and strong Linux skills. The position comes with a competitive salary and comprehensive benefits including flexible working and private medical insurance.
Senior Cloud Network Engineer
Cerebras Bristol, Gloucestershire
About Graphcore How often do you get the chance to build a technology that transforms the future of humanity? Graphcore products have set the standard in made-for-AI compute hardware and software, gaining global attention and industry acclaim. Now we are developing the next generation of artificial intelligence compute with systems that will allow AI researchers to develop more advanced models, help scientists unlock exciting new discoveries, and power companies around the world as they put AI at the heart of their business. Graphcore recently joined SoftBank Group - bringing large and ongoing investment from one of the world's leading backers of innovative AI companies. Job Summary We are looking for a Senior Network Engineer to join our Cloud Platform Team and help develop and deploy clouds and services. Working closely with our colleagues in Software Platform, Datacentre Operations and Product Development teams, you will deploy services on our fleet of cutting edge AI systems. As part of our Software Platform organisation, you will be involved in the cloud integration, validation, performance benchmarking, optimisation, and development of our high-performance AI solutions. These include in-house AI systems alongside off the shelf high-performance servers, switches and storage solutions. This is a hands on technical role requiring a solid background in the use of cloud infrastructure, deployment using Infrastructure as Code, observability, high-performance networking and storage systems. You may have been working in an IT organisation, a datacentre, a cloud provider or as a developer of orchestration or cloud services. The Software Platform team at Graphcore We build Graphcore products into large-scale AI solutions for our customers and the Cloud Platform Team is responsible for providing such systems to both internal users via private clouds and customers via our own public clouds. Often the internal systems will be using and developing pre-release hardware and software, so it's vital you are comfortable with unproven components. Responsibilities and Duties Develop and operate high-performance ethernet infrastructure on our private clouds and support internal users in their use. You will turn end user and product requirements into deployed services. Help to build automation to collect and analyse metrics and other data from the network infrastructure to support clear identification and reporting of any issues. Work with users to provide information of any product related issues to Engineering and QA departments. Work with our Datacentre Operations Engineers to maintain, tune and operate the fleet of AI systems at peak performance in our private clouds. Work with external vendors of off the shelf switches, servers and storage solutions to integrate third party products into our Cloud Reference Design, with a focus on network performance, automation and resilience. Skills and Experience (all required) Bachelor's degree or equivalent practical experience in a relevant subject. Significant hands on experience with one or more vendor's high end (100Gb/s+) Ethernet switch solutions. Experience managing on premises or private cloud environments. Solid software engineering or IT experience with a proven track record of delivering technical output as an individual contributor. Experience working in an AGILE and SCRUM framework, including understanding of priorities, risks, issues, impacts and constraints. Strong proven Linux scripting ability (bash, python, awk, sed). Strong proven Linux system administration (Ubuntu, RHEL and variants). Experience with a version control system (preferably Git) and using it to manage system configuration or automation. Experience with Continuous Integration or testing pipelines using GitLab, GitHub or similar. A solid hands on understanding of the technologies underpinning cloud services (APIs, virtualization of CPUs, IO, systems) and how they relate to high-performance networking. Experience with IAC automation tools (Terraform/OpenTofu, Ansible). Experience with container deployment and management tools (e.g. docker). Experience with solutions for monitoring and observability e.g. Grafana, Prometheus, OpenSearch/ElasticSearch, Loki. Good communication and presentation skills, and experience dealing with end users of IT services. An ability to work independently on critical infrastructure with minimal oversight, and with a focus on end user availability. Desirable but not required Experience with OpenStack cloud platforms. Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions. Experience with hardware offloading on RDMA capable NICs and how that integrates with virtual networking on Open vSwitch, KVM/QEMU. Experience with managing production Kubernetes clusters and workloads with an automation tool such as ArgoCD. Benefits In addition to a competitive salary, Graphcore offers flexible working, a generous annual leave policy, private medical insurance and health cash plan, a dental plan, pension (matched up to 5%), life assurance and income protection. We have a generous parental leave policy and an employee assistance programme (which includes health, mental wellbeing, and bereavement support). We offer a range of healthy food and snacks at our central Bristol office and have our own barista bar! We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments. Sponsorship Applicants for this position must hold the right to work in the UK. Unfortunately at this time, we are unable to provide visa sponsorship or support for visa applications.
11/07/2026
Full time
About Graphcore How often do you get the chance to build a technology that transforms the future of humanity? Graphcore products have set the standard in made-for-AI compute hardware and software, gaining global attention and industry acclaim. Now we are developing the next generation of artificial intelligence compute with systems that will allow AI researchers to develop more advanced models, help scientists unlock exciting new discoveries, and power companies around the world as they put AI at the heart of their business. Graphcore recently joined SoftBank Group - bringing large and ongoing investment from one of the world's leading backers of innovative AI companies. Job Summary We are looking for a Senior Network Engineer to join our Cloud Platform Team and help develop and deploy clouds and services. Working closely with our colleagues in Software Platform, Datacentre Operations and Product Development teams, you will deploy services on our fleet of cutting edge AI systems. As part of our Software Platform organisation, you will be involved in the cloud integration, validation, performance benchmarking, optimisation, and development of our high-performance AI solutions. These include in-house AI systems alongside off the shelf high-performance servers, switches and storage solutions. This is a hands on technical role requiring a solid background in the use of cloud infrastructure, deployment using Infrastructure as Code, observability, high-performance networking and storage systems. You may have been working in an IT organisation, a datacentre, a cloud provider or as a developer of orchestration or cloud services. The Software Platform team at Graphcore We build Graphcore products into large-scale AI solutions for our customers and the Cloud Platform Team is responsible for providing such systems to both internal users via private clouds and customers via our own public clouds. Often the internal systems will be using and developing pre-release hardware and software, so it's vital you are comfortable with unproven components. Responsibilities and Duties Develop and operate high-performance ethernet infrastructure on our private clouds and support internal users in their use. You will turn end user and product requirements into deployed services. Help to build automation to collect and analyse metrics and other data from the network infrastructure to support clear identification and reporting of any issues. Work with users to provide information of any product related issues to Engineering and QA departments. Work with our Datacentre Operations Engineers to maintain, tune and operate the fleet of AI systems at peak performance in our private clouds. Work with external vendors of off the shelf switches, servers and storage solutions to integrate third party products into our Cloud Reference Design, with a focus on network performance, automation and resilience. Skills and Experience (all required) Bachelor's degree or equivalent practical experience in a relevant subject. Significant hands on experience with one or more vendor's high end (100Gb/s+) Ethernet switch solutions. Experience managing on premises or private cloud environments. Solid software engineering or IT experience with a proven track record of delivering technical output as an individual contributor. Experience working in an AGILE and SCRUM framework, including understanding of priorities, risks, issues, impacts and constraints. Strong proven Linux scripting ability (bash, python, awk, sed). Strong proven Linux system administration (Ubuntu, RHEL and variants). Experience with a version control system (preferably Git) and using it to manage system configuration or automation. Experience with Continuous Integration or testing pipelines using GitLab, GitHub or similar. A solid hands on understanding of the technologies underpinning cloud services (APIs, virtualization of CPUs, IO, systems) and how they relate to high-performance networking. Experience with IAC automation tools (Terraform/OpenTofu, Ansible). Experience with container deployment and management tools (e.g. docker). Experience with solutions for monitoring and observability e.g. Grafana, Prometheus, OpenSearch/ElasticSearch, Loki. Good communication and presentation skills, and experience dealing with end users of IT services. An ability to work independently on critical infrastructure with minimal oversight, and with a focus on end user availability. Desirable but not required Experience with OpenStack cloud platforms. Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions. Experience with hardware offloading on RDMA capable NICs and how that integrates with virtual networking on Open vSwitch, KVM/QEMU. Experience with managing production Kubernetes clusters and workloads with an automation tool such as ArgoCD. Benefits In addition to a competitive salary, Graphcore offers flexible working, a generous annual leave policy, private medical insurance and health cash plan, a dental plan, pension (matched up to 5%), life assurance and income protection. We have a generous parental leave policy and an employee assistance programme (which includes health, mental wellbeing, and bereavement support). We offer a range of healthy food and snacks at our central Bristol office and have our own barista bar! We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments. Sponsorship Applicants for this position must hold the right to work in the UK. Unfortunately at this time, we are unable to provide visa sponsorship or support for visa applications.
HPC & Linux Platform Engineer - Self-Service, Automation
United States Digital Space LLC Milton Keynes, Buckinghamshire
United States Digital Space LLC is seeking an experienced Platform Engineer to design, build and operate scalable platform services for engineering and HPC workloads across hybrid cloud and on prem environments. You will implement self service capabilities, drive IaC, and ensure reliability and security across both on prem and cloud platforms. You will collaborate with software engineers to meet application requirements, maintain platform governance, and keep documentation up to date while
10/07/2026
Full time
United States Digital Space LLC is seeking an experienced Platform Engineer to design, build and operate scalable platform services for engineering and HPC workloads across hybrid cloud and on prem environments. You will implement self service capabilities, drive IaC, and ensure reliability and security across both on prem and cloud platforms. You will collaborate with software engineers to meet application requirements, maintain platform governance, and keep documentation up to date while
IT Platform Engineer - HPC & Linux
United States Digital Space LLC Milton Keynes, Buckinghamshire
Purpose Design, build, operate and continuously improve secure, scalable and reliable platform services supporting engineering, simulation and HPC workloads across hybrid cloud and on-premises environments. The role combines platform engineering, automation, infrastructure operations and developer enablement, delivering self-service capabilities and platform products that improve the productivity, reliability and efficiency of engineering teams. Accountabilities Design, implement and maintain innovative platforms end-to-end for the Technology Campus Design and develop self-service platform capabilities that enable engineering teams to provision and consume infrastructure, compute, storage and application services consistently and securely. Define and maintain service reliability objectives, capacity plans and operational metrics to ensure platform availability, performance and scalability. Provide advanced technical support and manage problem escalations from users communicating well and recording in ticketing management system. Partner with software engineering teams to understand application requirements, improve developer experience, support cloud-native delivery practices and ensure platform services effectively meet the needs of engineering workloads. Own on-premises and cloud-hosted platform services, proactively identifying reliability, performance and scalability improvements and leading the design and implementation of appropriate solutions. Drive automation and Infrastructure as Code practices to improve reliability, consistency, and operational efficiency of platform services Contribute to platform governance forums, providing technical input, sharing knowledge, and supporting alignment across infrastructure and engineering teams Enable efficient and reliable consumption of platform services by internal teams, with a focus on usability, repeatability, and reduced operational friction Additional Accountabilities Participate in wider team projects, change management and take an active role in reviewing architecture of new solutions across Platforms infrastructure and HPC Contribute to platform architecture, technology roadmaps and lifecycle management decisions to ensure services remain fit for purpose and aligned with future business requirements. Maintain accurate platform documentation, asset records, and operational runbooks to support effective operation and knowledge sharing Ensure platform solutions comply with organisational security standards, policies, and regulatory requirements Work closely with vendors to leverage their expertise and solutions, making use of technical partnerships to improve performance, inform future technical direction, execute proof of concepts and to implement new technology Essential Competencies Expert knowledge administering Linux/Unix systems Experience in scripting and programming e.g. Python, Bash, Go Experience implementing CI/CD and GitOps workflows using tools such as GitHub Actions, GitLab, ArgoCD or Flux. Working knowledge of platform security principles including identity management, secrets management, vulnerability remediation, hardening and least-privilege access controls. Network knowledge of InfiniBand, MPI and Ethernet concepts Working knowledge of platform security principles. Experience with configuration management and infrastructure as code tools (Ansible, Git, Vault) Knowledge of container orchestration tools, virtualisation and observability stacks e.g. Kubernetes, Grafana, Kafka, Docker, OpenStack, OLVM. Experience implementing platform and application observability solutions including monitoring, logging, tracing, alerting and telemetry. Awareness of secure software development and DevSecOps principles. Experience working within Agile and DevOps delivery models. Knowledge of software delivery platforms such as GitHub, GitLab, Azure DevOps or equivalent. Experience designing, deploying and operating infrastructure services within public cloud platforms.
10/07/2026
Full time
Purpose Design, build, operate and continuously improve secure, scalable and reliable platform services supporting engineering, simulation and HPC workloads across hybrid cloud and on-premises environments. The role combines platform engineering, automation, infrastructure operations and developer enablement, delivering self-service capabilities and platform products that improve the productivity, reliability and efficiency of engineering teams. Accountabilities Design, implement and maintain innovative platforms end-to-end for the Technology Campus Design and develop self-service platform capabilities that enable engineering teams to provision and consume infrastructure, compute, storage and application services consistently and securely. Define and maintain service reliability objectives, capacity plans and operational metrics to ensure platform availability, performance and scalability. Provide advanced technical support and manage problem escalations from users communicating well and recording in ticketing management system. Partner with software engineering teams to understand application requirements, improve developer experience, support cloud-native delivery practices and ensure platform services effectively meet the needs of engineering workloads. Own on-premises and cloud-hosted platform services, proactively identifying reliability, performance and scalability improvements and leading the design and implementation of appropriate solutions. Drive automation and Infrastructure as Code practices to improve reliability, consistency, and operational efficiency of platform services Contribute to platform governance forums, providing technical input, sharing knowledge, and supporting alignment across infrastructure and engineering teams Enable efficient and reliable consumption of platform services by internal teams, with a focus on usability, repeatability, and reduced operational friction Additional Accountabilities Participate in wider team projects, change management and take an active role in reviewing architecture of new solutions across Platforms infrastructure and HPC Contribute to platform architecture, technology roadmaps and lifecycle management decisions to ensure services remain fit for purpose and aligned with future business requirements. Maintain accurate platform documentation, asset records, and operational runbooks to support effective operation and knowledge sharing Ensure platform solutions comply with organisational security standards, policies, and regulatory requirements Work closely with vendors to leverage their expertise and solutions, making use of technical partnerships to improve performance, inform future technical direction, execute proof of concepts and to implement new technology Essential Competencies Expert knowledge administering Linux/Unix systems Experience in scripting and programming e.g. Python, Bash, Go Experience implementing CI/CD and GitOps workflows using tools such as GitHub Actions, GitLab, ArgoCD or Flux. Working knowledge of platform security principles including identity management, secrets management, vulnerability remediation, hardening and least-privilege access controls. Network knowledge of InfiniBand, MPI and Ethernet concepts Working knowledge of platform security principles. Experience with configuration management and infrastructure as code tools (Ansible, Git, Vault) Knowledge of container orchestration tools, virtualisation and observability stacks e.g. Kubernetes, Grafana, Kafka, Docker, OpenStack, OLVM. Experience implementing platform and application observability solutions including monitoring, logging, tracing, alerting and telemetry. Awareness of secure software development and DevSecOps principles. Experience working within Agile and DevOps delivery models. Knowledge of software delivery platforms such as GitHub, GitLab, Azure DevOps or equivalent. Experience designing, deploying and operating infrastructure services within public cloud platforms.
Senior Cloud Engineer K8S
Cerebras Bristol, Gloucestershire
About Graphcore How often do you get the chance to build a technology that transforms the future of humanity? Graphcore products have set the standard in made-for-AI compute hardware and software, gaining global attention and industry acclaim. Now we are developing the next generation of artificial intelligence compute with systems that will allow AI researchers to develop more advanced models, help scientists unlock exciting new discoveries, and power companies around the world as they put AI at the heart of their business. Graphcore recently joined SoftBank Group - bringing large and ongoing investment from one of the world's leading backers of innovative AI companies. Job Summary We are looking for a Senior Engineer to join our Cloud Platform Team and help develop and deploy clouds and services. Working closely with our colleagues in Software Platform, Datacentre Operations and Product Development teams, you will deploy services on our fleet of cutting edge AI systems. As part of our Software Platform organisation, you will be involved in the cloud integration, validation, performance benchmarking, optimisation, and development of our high-performance AI solutions. These include in house AI systems alongside off the shelf high-performance servers, switches and storage solutions. This is a hand on technical role requiring a solid background in the use of cloud infrastructure, deployment using Infrastructure as Code, observability, high-performance networking and storage systems. You may have been working in an IT organisation, a datacentre, a cloud provider or as a developer of orchestration or cloud services. The Software Platform team at Graphcore We build Graphcore products into large scale AI solutions for our customers and the Cloud Platform Team is responsible for providing such systems to both internal users via private clouds and customers via our own public clouds. Often the internal systems will be using and developing pre release hardware and software, so it's vital you are comfortable with unproven components. Responsibilities and Duties Develop and operate Kubernetes managed end user services on our private clouds and support internal users in their use. You will turn end user and product requirements into deployed services. Work with our Datacentre Operations Engineers to maintain and operate the fleet of AI systems at peak performance in our private clouds. Configure and test new Graphcore AI hardware and systems using Continuous Deployment and Infrastructure as code in internal and external datacentres. Skills and Experience (all required) Bachelor's degree or equivalent practical experience in a relevant subject. Experience with managing production Kubernetes clusters and workloads with a continuous delivery tool such as ArgoCD. Solid software engineering or IT experience with a proven track record of delivering technical output as an individual contributor. Experience working in an AGILE and SCRUM framework, including understanding of priorities, risks, issues, impacts and constraints. Strong proven Linux scripting ability (bash, python, awk, sed). Strong proven Linux system administration (Ubuntu, RHEL and variants). Experience with a version control system (preferably Git) and using it to manage system configuration or automation. Experience with Continuous Integration or testing pipelines using GitLab, GitHub or similar. A solid hands on understanding of the technologies underpinning cloud services (APIs, virtualisation of CPUs, IO, systems), virtual networks, block storage, resource management and monitoring. Experience with IAC automation tools (Terraform/OpenTofu, Ansible, Packer). Good communication and presentation skills, and experience dealing with end users of IT services. An ability to work independently on critical infrastructure with minimal oversight, and with a focus on end user availability. Desirable but not required: Experience with OpenStack cloud platform(s). Experience with solutions for monitoring and observability. e.g. Grafana, Prometheus, OpenSearch/ElasticSearch, Loki. Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions. Programming experience with Python3 utilising classes and inheritance. Benefits In addition to a competitive salary, Graphcore offers flexible working, a generous annual leave policy, private medical insurance and health cash plan, a dental plan, pension (matched up to 5%), life assurance and income protection. We have a generous parental leave policy and an employee assistance programme (which includes health, mental wellbeing, and bereavement support). We offer a range of healthy food and snacks at our central Bristol office and have our own barista bar! We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments. Sponsorship Applicants for this position must hold the right to work in the UK. Unfortunately at this time, we are unable to provide visa sponsorship or support for visa applications.
09/07/2026
Full time
About Graphcore How often do you get the chance to build a technology that transforms the future of humanity? Graphcore products have set the standard in made-for-AI compute hardware and software, gaining global attention and industry acclaim. Now we are developing the next generation of artificial intelligence compute with systems that will allow AI researchers to develop more advanced models, help scientists unlock exciting new discoveries, and power companies around the world as they put AI at the heart of their business. Graphcore recently joined SoftBank Group - bringing large and ongoing investment from one of the world's leading backers of innovative AI companies. Job Summary We are looking for a Senior Engineer to join our Cloud Platform Team and help develop and deploy clouds and services. Working closely with our colleagues in Software Platform, Datacentre Operations and Product Development teams, you will deploy services on our fleet of cutting edge AI systems. As part of our Software Platform organisation, you will be involved in the cloud integration, validation, performance benchmarking, optimisation, and development of our high-performance AI solutions. These include in house AI systems alongside off the shelf high-performance servers, switches and storage solutions. This is a hand on technical role requiring a solid background in the use of cloud infrastructure, deployment using Infrastructure as Code, observability, high-performance networking and storage systems. You may have been working in an IT organisation, a datacentre, a cloud provider or as a developer of orchestration or cloud services. The Software Platform team at Graphcore We build Graphcore products into large scale AI solutions for our customers and the Cloud Platform Team is responsible for providing such systems to both internal users via private clouds and customers via our own public clouds. Often the internal systems will be using and developing pre release hardware and software, so it's vital you are comfortable with unproven components. Responsibilities and Duties Develop and operate Kubernetes managed end user services on our private clouds and support internal users in their use. You will turn end user and product requirements into deployed services. Work with our Datacentre Operations Engineers to maintain and operate the fleet of AI systems at peak performance in our private clouds. Configure and test new Graphcore AI hardware and systems using Continuous Deployment and Infrastructure as code in internal and external datacentres. Skills and Experience (all required) Bachelor's degree or equivalent practical experience in a relevant subject. Experience with managing production Kubernetes clusters and workloads with a continuous delivery tool such as ArgoCD. Solid software engineering or IT experience with a proven track record of delivering technical output as an individual contributor. Experience working in an AGILE and SCRUM framework, including understanding of priorities, risks, issues, impacts and constraints. Strong proven Linux scripting ability (bash, python, awk, sed). Strong proven Linux system administration (Ubuntu, RHEL and variants). Experience with a version control system (preferably Git) and using it to manage system configuration or automation. Experience with Continuous Integration or testing pipelines using GitLab, GitHub or similar. A solid hands on understanding of the technologies underpinning cloud services (APIs, virtualisation of CPUs, IO, systems), virtual networks, block storage, resource management and monitoring. Experience with IAC automation tools (Terraform/OpenTofu, Ansible, Packer). Good communication and presentation skills, and experience dealing with end users of IT services. An ability to work independently on critical infrastructure with minimal oversight, and with a focus on end user availability. Desirable but not required: Experience with OpenStack cloud platform(s). Experience with solutions for monitoring and observability. e.g. Grafana, Prometheus, OpenSearch/ElasticSearch, Loki. Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions. Programming experience with Python3 utilising classes and inheritance. Benefits In addition to a competitive salary, Graphcore offers flexible working, a generous annual leave policy, private medical insurance and health cash plan, a dental plan, pension (matched up to 5%), life assurance and income protection. We have a generous parental leave policy and an employee assistance programme (which includes health, mental wellbeing, and bereavement support). We offer a range of healthy food and snacks at our central Bristol office and have our own barista bar! We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments. Sponsorship Applicants for this position must hold the right to work in the UK. Unfortunately at this time, we are unable to provide visa sponsorship or support for visa applications.
Technical Futures Ltd
Platform Engineer
Technical Futures Ltd Cambridge, Cambridgeshire
Exciting deeptech start-up seeks a bright, forward thinking Platform Engineer to join their dedicated team. With strong software engineering skills in modern languages ( ideally to include Python and Rust) backed a good academic history, you'll bring curiosity about scientific computing and care for code quality. Applications are welcomed from mid level up to Senior level Engineers with knowledge of Cloud computing or HPC job management (such as Slurm), identity and authorization flows (such as OAuth2/OIDC) being highly beneficial. This cutting edge technology company, focused on optimizing complex engineering systems, seeks a top class Platform Engineer who will confidently work collaboratively with engineers and scientists to have a real say in how the platform layer is designed. You will ideally bring some of the following skills and experience: Strong academic background. Could be Software Engineering, Computer Science, Physics or Scientific Software related. Strong software engineering skills in modern languages (Python & Rust ideal). Experience designing systems and the APIs between them - services that coordinate work, manage state and handle failure. Focus on code quality. Some of the following should compliment the skills above: Experience of Cloud computing or HPC job management ( such as Slurm). Identity and authorization flows such as OIDC/OAuth2. Deploying containerized services on Linux (such as Podman). Infrastructure as code (such as Ansible, OpenTofu). Modern Python tooling and packaging. Data storage and pipelines, including large simulation results and model serving. Hybrid working available (3 days office /2 WFH), a very generous base salary dependent on your level of skills and experience and benefits to include Shares, 30 days holiday + time off between Xmas and New Year, Private healthcare, Pension Plan, Life Assurance and much more.
09/07/2026
Full time
Exciting deeptech start-up seeks a bright, forward thinking Platform Engineer to join their dedicated team. With strong software engineering skills in modern languages ( ideally to include Python and Rust) backed a good academic history, you'll bring curiosity about scientific computing and care for code quality. Applications are welcomed from mid level up to Senior level Engineers with knowledge of Cloud computing or HPC job management (such as Slurm), identity and authorization flows (such as OAuth2/OIDC) being highly beneficial. This cutting edge technology company, focused on optimizing complex engineering systems, seeks a top class Platform Engineer who will confidently work collaboratively with engineers and scientists to have a real say in how the platform layer is designed. You will ideally bring some of the following skills and experience: Strong academic background. Could be Software Engineering, Computer Science, Physics or Scientific Software related. Strong software engineering skills in modern languages (Python & Rust ideal). Experience designing systems and the APIs between them - services that coordinate work, manage state and handle failure. Focus on code quality. Some of the following should compliment the skills above: Experience of Cloud computing or HPC job management ( such as Slurm). Identity and authorization flows such as OIDC/OAuth2. Deploying containerized services on Linux (such as Podman). Infrastructure as code (such as Ansible, OpenTofu). Modern Python tooling and packaging. Data storage and pipelines, including large simulation results and model serving. Hybrid working available (3 days office /2 WFH), a very generous base salary dependent on your level of skills and experience and benefits to include Shares, 30 days holiday + time off between Xmas and New Year, Private healthcare, Pension Plan, Life Assurance and much more.
Principal Machine Learning Infrastructure Engineer London, United Kingdom
PhysicsX Ltd
Senior Machine Learning Infrastructure Engineer London, United Kingdom About us PhysicsX is a deep-tech company with roots in numerical physics and Formula One, dedicated to accelerating hardware innovation at the speed of software. We are building an AI-driven simulation software stack for engineering and manufacturing across advanced industries. By enabling high-fidelity, multi-physics simulation through AI inference across the entire engineering lifecycle, PhysicsX unlocks new levels of optimization and automation in design, manufacturing, and operations - empowering engineers to push the boundaries of possibility. Our customers include leading innovators in Aerospace & Defense, Materials, Energy, Semiconductors, and Automotive. Note:We are currently recruiting for multiple positions, however please only apply for the role that best aligns with your skillset and career goals. The Role The Senior ML Infrastructure Engineer will extend and operate the infrastructure that powers our research model training, fine-tuning, and serving pipelines. You will be embedded within our Research function, partnering directly with ML engineers and research scientists to ensure they can train Large Physics Models efficiently and reliably at scale. Team Context In this role, you will be vertically embedded in Research, working daily with: Research Scientists who determine the model architectures and methods ML Engineers who implement and develop the models Simulation Data Engineers who are accountable for upstream data pipelines You will have end-to-end responsibilities over the research infrastructure, with the autonomy to make architectural decisions and the responsibility to keep data flowing reliably. Horizontally, you will be part of an infrastructure engineering group responsible for infrastructure across the company. What you will do Training Infrastructure Design and operate distributed training infrastructure for neural operator architectures (Transolver, Point Cloud Transformer, etc.) on our large NVIDIA DGX B200 platform. Optimize training pipelines for throughput, fault tolerance, and cost efficiency, including checkpointing strategies, gradient accumulation, and multi-node synchronization. Build and maintain experiment tracking and observability systems that give researchers clear visibility into training runs, hyperparameter sweeps, and model performance. Data I/O and Performance Solve data loading bottlenecks for large-scale mesh datasets. Optimize data pipelines for efficient I/O from cloud storage, including prefetching, caching, and format optimization. Work with heterogeneous data sources of varying formats and resolutions. Model Serving and Deployment Build serving infrastructure for pre-trained LPMs, supporting both zero shot inference and uncertainty quantification (Monte Carlo Dropout). Design and implement model packaging pipelines for customer deployment. Models must run reliably in customer environments with fine tuning capabilities. Ensure reproducibility: any model checkpoint should be deployable with consistent behaviour. Platform and Tooling Improve developer experience for the Research team with fast iteration cycles, reliable CI/CD, clear debugging tools. Collaborate with the broader Infrastructure team on shared patterns and standards. What you bring to the table Ability to scope and effectively deliver projects, prioritising activity as needed. Problem solving skills and the ability to analyse issues, identify causes, and recommend solutions quickly. Excellent collaboration and communication skills, especially in a research setting. You can translate "the model isn't converging" into infrastructure hypotheses and solutions, and can bridge technical abstractions with implementations. 5+ years of experience building and operating ML infrastructure at scale: Deep expertise in distributed training: you've debugged NCCL hangs, optimized collective communication, and know when to use FSDP vs. DDP vs. pipeline parallelism Strong systems fundamentals: Linux, networking (including domain specific NVLink and InfiniBand), storage I/O, profiling and performance optimization Production experience with Kubernetes and SLURM for job orchestration on GPU clusters Proficiency in Python and ML frameworks (PyTorch strongly preferred) Experience with cloud GPU infrastructure; ideally CoreWeave or similar GPU/HPC-focused clouds Ideally Experience with geometric deep learning or neural operators, architectures that operate on meshes, point clouds, or graphs Background in HPC for simulation engineering, familiarity with how CFD/FEA workflows generate and consume data Experience building model serving infrastructure with latency and throughput requirements Familiarity with experiment tracking tools (Weights & Biases, MLflow) and observability stacks (Prometheus, Grafana) What we offer Equity options - share in our success and growth. 10% employer pension contribution - invest in your future. Free office lunches - great food to fuel your workdays. Flexible working - balance your work and life in a way that works for you. Hybrid setup - enjoy our new Shoreditch office while keeping remote flexibility. Enhanced parental leave - support for life's biggest milestones. Private healthcare - comprehensive coverage Personal development - access learning and training to help you grow. Work from anywhere - extend your remote setup to enjoy the sun or reconnect with loved ones. We value diversity and are committed to equal employment opportunity regardless of sex, race, religion, ethnicity, nationality, disability, age, sexual orientation or gender identity. We strongly encourage individuals from groups traditionally underrepresented in tech to apply. To help make a change, we sponsor bright women from disadvantaged backgrounds through their university degrees in science and mathematics. We collect diversity and inclusion data solely for the purpose of monitoring the effectiveness of our equal opportunities policies and ensuring compliance with UK employment and equality legislation. This information is confidential, used only in aggregate form, and will not influence the outcome of your application.
08/07/2026
Full time
Senior Machine Learning Infrastructure Engineer London, United Kingdom About us PhysicsX is a deep-tech company with roots in numerical physics and Formula One, dedicated to accelerating hardware innovation at the speed of software. We are building an AI-driven simulation software stack for engineering and manufacturing across advanced industries. By enabling high-fidelity, multi-physics simulation through AI inference across the entire engineering lifecycle, PhysicsX unlocks new levels of optimization and automation in design, manufacturing, and operations - empowering engineers to push the boundaries of possibility. Our customers include leading innovators in Aerospace & Defense, Materials, Energy, Semiconductors, and Automotive. Note:We are currently recruiting for multiple positions, however please only apply for the role that best aligns with your skillset and career goals. The Role The Senior ML Infrastructure Engineer will extend and operate the infrastructure that powers our research model training, fine-tuning, and serving pipelines. You will be embedded within our Research function, partnering directly with ML engineers and research scientists to ensure they can train Large Physics Models efficiently and reliably at scale. Team Context In this role, you will be vertically embedded in Research, working daily with: Research Scientists who determine the model architectures and methods ML Engineers who implement and develop the models Simulation Data Engineers who are accountable for upstream data pipelines You will have end-to-end responsibilities over the research infrastructure, with the autonomy to make architectural decisions and the responsibility to keep data flowing reliably. Horizontally, you will be part of an infrastructure engineering group responsible for infrastructure across the company. What you will do Training Infrastructure Design and operate distributed training infrastructure for neural operator architectures (Transolver, Point Cloud Transformer, etc.) on our large NVIDIA DGX B200 platform. Optimize training pipelines for throughput, fault tolerance, and cost efficiency, including checkpointing strategies, gradient accumulation, and multi-node synchronization. Build and maintain experiment tracking and observability systems that give researchers clear visibility into training runs, hyperparameter sweeps, and model performance. Data I/O and Performance Solve data loading bottlenecks for large-scale mesh datasets. Optimize data pipelines for efficient I/O from cloud storage, including prefetching, caching, and format optimization. Work with heterogeneous data sources of varying formats and resolutions. Model Serving and Deployment Build serving infrastructure for pre-trained LPMs, supporting both zero shot inference and uncertainty quantification (Monte Carlo Dropout). Design and implement model packaging pipelines for customer deployment. Models must run reliably in customer environments with fine tuning capabilities. Ensure reproducibility: any model checkpoint should be deployable with consistent behaviour. Platform and Tooling Improve developer experience for the Research team with fast iteration cycles, reliable CI/CD, clear debugging tools. Collaborate with the broader Infrastructure team on shared patterns and standards. What you bring to the table Ability to scope and effectively deliver projects, prioritising activity as needed. Problem solving skills and the ability to analyse issues, identify causes, and recommend solutions quickly. Excellent collaboration and communication skills, especially in a research setting. You can translate "the model isn't converging" into infrastructure hypotheses and solutions, and can bridge technical abstractions with implementations. 5+ years of experience building and operating ML infrastructure at scale: Deep expertise in distributed training: you've debugged NCCL hangs, optimized collective communication, and know when to use FSDP vs. DDP vs. pipeline parallelism Strong systems fundamentals: Linux, networking (including domain specific NVLink and InfiniBand), storage I/O, profiling and performance optimization Production experience with Kubernetes and SLURM for job orchestration on GPU clusters Proficiency in Python and ML frameworks (PyTorch strongly preferred) Experience with cloud GPU infrastructure; ideally CoreWeave or similar GPU/HPC-focused clouds Ideally Experience with geometric deep learning or neural operators, architectures that operate on meshes, point clouds, or graphs Background in HPC for simulation engineering, familiarity with how CFD/FEA workflows generate and consume data Experience building model serving infrastructure with latency and throughput requirements Familiarity with experiment tracking tools (Weights & Biases, MLflow) and observability stacks (Prometheus, Grafana) What we offer Equity options - share in our success and growth. 10% employer pension contribution - invest in your future. Free office lunches - great food to fuel your workdays. Flexible working - balance your work and life in a way that works for you. Hybrid setup - enjoy our new Shoreditch office while keeping remote flexibility. Enhanced parental leave - support for life's biggest milestones. Private healthcare - comprehensive coverage Personal development - access learning and training to help you grow. Work from anywhere - extend your remote setup to enjoy the sun or reconnect with loved ones. We value diversity and are committed to equal employment opportunity regardless of sex, race, religion, ethnicity, nationality, disability, age, sexual orientation or gender identity. We strongly encourage individuals from groups traditionally underrepresented in tech to apply. To help make a change, we sponsor bright women from disadvantaged backgrounds through their university degrees in science and mathematics. We collect diversity and inclusion data solely for the purpose of monitoring the effectiveness of our equal opportunities policies and ensuring compliance with UK employment and equality legislation. This information is confidential, used only in aggregate form, and will not influence the outcome of your application.
Senior Simulation Data Engineer London, United Kingdom
PhysicsX Ltd
PhysicsX is a deep-tech company with roots in numerical physics and Formula One, dedicated to accelerating hardware innovation at the speed of software. We are building an AI-driven simulation software stack for engineering and manufacturing across advanced industries. By enabling high-fidelity, multi-physics simulation through AI inference across the entire engineering lifecycle, PhysicsX unlocks new levels of optimization and automation in design, manufacturing, and operations - empowering engineers to push the boundaries of possibility. Our customers include leading innovators in Aerospace & Defense, Materials, Energy, Semiconductors, and Automotive. Note:We are currently recruiting for multiple positions, however please only apply for the role that best aligns with your skillset and career goals. The Role The Senior Simulation Data Engineer will extend and operate the infrastructure that powers our research Data Factory. You will be responsible for the end-to-end pipeline: from geometry preparation and simulation orchestration through validation, post-processing, and delivery to downstream ML training systems, using PhysicsX platform orchestration services where synergies exist. This role sits at the intersection of HPC engineering and data engineering. You will orchestrate long-running CFD simulations at scale, build robust data pipelines, and ensure that every simulation we produce meets rigorous quality standards. Team Context In this role, you will be vertically embedded in Research , working daily with: Research Scientists who define data requirements and quality standards ML Engineers who consume Data Factory outputs for model training ML Infrastructure Engineers who are accountable for downstream training infrastructure You will have end-to-end responsibilities over the Data Factory, with the autonomy to make architectural decisions and the responsibility to keep data flowing reliably. Horizontally, you will be part of an infrastructure engineering group responsible for infrastructure across the company. What you will do Simulation Orchestration Extend and operate the Data Factory infrastructure that orchestrates thousands of CFD simulations per day on cloud compute Design and operate job scheduling systems that maximize throughput while handling failures gracefully Build monitoring and alerting to detect simulation failures, convergence issues, and resource bottlenecks early Build high-performance data pipelines that move simulation outputs from solver results to ML-ready training data Implement geometry preprocessing workflows (mesh preparation, morphing, watertightness validation) Design and operate post-processing pipelines: surface decimation, field interpolation, format conversion Optimize I/O performance for large mesh datasets Data Quality and Validation Implement comprehensive validation checks at every pipeline stage: solver convergence, physical field bounds, post-processing fidelity Build systems that capture and quarantine bad data before they reach training pipelines Track and report data quality metrics across the entire Data Factory Work towards full provenance: training samples should be traceable back to their source geometry and simulation configuration Integration and Delivery Deliver validated datasets to downstream ML training infrastructure in formats optimized for efficient data loading Design data versioning and cataloging systems that support reproducible training runs Work closely with ML Infrastructure Engineers to ensure smooth handoff between data production and model training Support multi-dataset training workflows What you bring to the table Ability to scope and effectively deliver projects, prioritising activity as needed. Problem solving skills and the ability to analyse issues, identify causes, and recommend solutions quickly. Excellent collaboration and communication skills, especially in a research setting. You can translate "the model isn't converging" into infrastructure hypotheses and solutions, and can bridge technical abstractions with implementations. 5+ years of experience in data engineering, HPC engineering, or simulation infrastructure. Strong experience with orchestration systems: SLURM, Kubernetes, Temporal Production data pipeline experience: you've built and operated pipelines that process large volumes of data reliably Proficiency in Python for pipeline development and automation Systems engineering fundamentals: Linux, networking, storage systems, performance debugging Experience with cloud infrastructure; ideally CoreWeave or similar GPU/HPC focused clouds Background in HPC for simulation engineering: experience with CFD, FEA, or similar computational workflows (StarCCM+, OpenFOAM, ANSYS, etc.) Experience with geometry processing: mesh manipulation, CAD formats, PyVista Familiarity with scientific data formats: HDF5, VTK, NetCDF, Zarr Data quality engineering experience: validation frameworks, anomaly detection, data observability Ideally Understanding of CFD fundamentals, enough to interpret solver outputs and validation metrics Experience with 3D geometry pipelines (mesh decimation, field interpolation) Familiarity with ML data loading patterns and how training systems consume data What we offer Equity options - share in our success and growth. 10% employer pension contribution - invest in your future. Free office lunches - great food to fuel your workdays. Flexible working - balance your work and life in a way that works for you. Hybrid setup - enjoy our new Shoreditch office while keeping remote flexibility. Enhanced parental leave - support for life's biggest milestones. Private healthcare - comprehensive coverage Personal development - access learning and training to help you grow. Work from anywhere - extend your remote setup to enjoy the sun or reconnect with loved ones. We value diversity and are committed to equal employment opportunity regardless of sex, race, religion, ethnicity, nationality, disability, age, sexual orientation or gender identity. We strongly encourage individuals from groups traditionally underrepresented in tech to apply. To help make a change, we sponsor bright women from disadvantaged backgrounds through their university degrees in science and mathematics. We collect diversity and inclusion data solely for the purpose of monitoring the effectiveness of our equal opportunities policies and ensuring compliance with UK employment and equality legislation. This information is confidential, used only in aggregate form, and will not influence the outcome of your application.
08/07/2026
Full time
PhysicsX is a deep-tech company with roots in numerical physics and Formula One, dedicated to accelerating hardware innovation at the speed of software. We are building an AI-driven simulation software stack for engineering and manufacturing across advanced industries. By enabling high-fidelity, multi-physics simulation through AI inference across the entire engineering lifecycle, PhysicsX unlocks new levels of optimization and automation in design, manufacturing, and operations - empowering engineers to push the boundaries of possibility. Our customers include leading innovators in Aerospace & Defense, Materials, Energy, Semiconductors, and Automotive. Note:We are currently recruiting for multiple positions, however please only apply for the role that best aligns with your skillset and career goals. The Role The Senior Simulation Data Engineer will extend and operate the infrastructure that powers our research Data Factory. You will be responsible for the end-to-end pipeline: from geometry preparation and simulation orchestration through validation, post-processing, and delivery to downstream ML training systems, using PhysicsX platform orchestration services where synergies exist. This role sits at the intersection of HPC engineering and data engineering. You will orchestrate long-running CFD simulations at scale, build robust data pipelines, and ensure that every simulation we produce meets rigorous quality standards. Team Context In this role, you will be vertically embedded in Research , working daily with: Research Scientists who define data requirements and quality standards ML Engineers who consume Data Factory outputs for model training ML Infrastructure Engineers who are accountable for downstream training infrastructure You will have end-to-end responsibilities over the Data Factory, with the autonomy to make architectural decisions and the responsibility to keep data flowing reliably. Horizontally, you will be part of an infrastructure engineering group responsible for infrastructure across the company. What you will do Simulation Orchestration Extend and operate the Data Factory infrastructure that orchestrates thousands of CFD simulations per day on cloud compute Design and operate job scheduling systems that maximize throughput while handling failures gracefully Build monitoring and alerting to detect simulation failures, convergence issues, and resource bottlenecks early Build high-performance data pipelines that move simulation outputs from solver results to ML-ready training data Implement geometry preprocessing workflows (mesh preparation, morphing, watertightness validation) Design and operate post-processing pipelines: surface decimation, field interpolation, format conversion Optimize I/O performance for large mesh datasets Data Quality and Validation Implement comprehensive validation checks at every pipeline stage: solver convergence, physical field bounds, post-processing fidelity Build systems that capture and quarantine bad data before they reach training pipelines Track and report data quality metrics across the entire Data Factory Work towards full provenance: training samples should be traceable back to their source geometry and simulation configuration Integration and Delivery Deliver validated datasets to downstream ML training infrastructure in formats optimized for efficient data loading Design data versioning and cataloging systems that support reproducible training runs Work closely with ML Infrastructure Engineers to ensure smooth handoff between data production and model training Support multi-dataset training workflows What you bring to the table Ability to scope and effectively deliver projects, prioritising activity as needed. Problem solving skills and the ability to analyse issues, identify causes, and recommend solutions quickly. Excellent collaboration and communication skills, especially in a research setting. You can translate "the model isn't converging" into infrastructure hypotheses and solutions, and can bridge technical abstractions with implementations. 5+ years of experience in data engineering, HPC engineering, or simulation infrastructure. Strong experience with orchestration systems: SLURM, Kubernetes, Temporal Production data pipeline experience: you've built and operated pipelines that process large volumes of data reliably Proficiency in Python for pipeline development and automation Systems engineering fundamentals: Linux, networking, storage systems, performance debugging Experience with cloud infrastructure; ideally CoreWeave or similar GPU/HPC focused clouds Background in HPC for simulation engineering: experience with CFD, FEA, or similar computational workflows (StarCCM+, OpenFOAM, ANSYS, etc.) Experience with geometry processing: mesh manipulation, CAD formats, PyVista Familiarity with scientific data formats: HDF5, VTK, NetCDF, Zarr Data quality engineering experience: validation frameworks, anomaly detection, data observability Ideally Understanding of CFD fundamentals, enough to interpret solver outputs and validation metrics Experience with 3D geometry pipelines (mesh decimation, field interpolation) Familiarity with ML data loading patterns and how training systems consume data What we offer Equity options - share in our success and growth. 10% employer pension contribution - invest in your future. Free office lunches - great food to fuel your workdays. Flexible working - balance your work and life in a way that works for you. Hybrid setup - enjoy our new Shoreditch office while keeping remote flexibility. Enhanced parental leave - support for life's biggest milestones. Private healthcare - comprehensive coverage Personal development - access learning and training to help you grow. Work from anywhere - extend your remote setup to enjoy the sun or reconnect with loved ones. We value diversity and are committed to equal employment opportunity regardless of sex, race, religion, ethnicity, nationality, disability, age, sexual orientation or gender identity. We strongly encourage individuals from groups traditionally underrepresented in tech to apply. To help make a change, we sponsor bright women from disadvantaged backgrounds through their university degrees in science and mathematics. We collect diversity and inclusion data solely for the purpose of monitoring the effectiveness of our equal opportunities policies and ensuring compliance with UK employment and equality legislation. This information is confidential, used only in aggregate form, and will not influence the outcome of your application.

Modal Window

  • Home
  • Contact
  • About Us
  • FAQs
  • Terms & Conditions
  • Privacy
  • Employer
  • Post a Job
  • Search Resumes
  • Sign in
  • Job Seeker
  • Find Jobs
  • Create Resume
  • Sign in
  • IT blog
  • Facebook
  • Twitter
  • LinkedIn
  • Youtube
© 2008-2026 IT Job Board