Role: Senior Site Reliability Engineer (SRE) - Kubernetes / OpenShift Location: Remote - UK (possible paid occasional travel to TIG Secure site locations as required) Job Type: Full-time, Permanent (37.5 hours) Salary: Competitive + benefits + package Security Clearance Requirements Please note that holding a current Security Clearance is not essential at the time of application, but eligibility is required. This role requires the successful candidate to be eligible for Security Check (SC) clearance. To meet this requirement, applicants must: Have the right to work in the UK Have lived in the UK continuously for the past 5 years Not have spent more than 6 months outside the UK in total during that period Be willing to undergo security vetting as part of the onboarding process About You You're an experienced SRE, Platform Engineer or Cloud Engineer with strong hands on experience running Kubernetes in production environments. You're comfortable working across Linux, Kubernetes, cloud native tooling, automation, observability, CI/CD and infrastructure as code. You understand that reliability, security and operational maturity are critical to how modern platforms support engineering teams and customer facing services. You enjoy treating infrastructure as a product, automating repeatable work, improving resilience, and building platforms that other engineers can rely on. You're calm under pressure, methodical during incidents, and able to turn operational challenges into long term improvements. You may have worked in a regulated, secure, government, defence, financial services, telecoms, managed services or cloud native environment, but most importantly you have operated Kubernetes at depth and understand the realities of production ownership. You're a senior individual contributor who can mentor others, influence engineering practice, and provide technical authority without needing formal line management responsibility. About the Role We're looking for a Senior Site Reliability Engineer (SRE) to help operate, harden and mature our production OKD / Kubernetes platforms. This is a hands on engineering role focused on reliability, automation, observability, GitOps, CI/CD and secure platform operations. You'll work across the full stack, from bare metal and virtualisation through to Kubernetes control plane operations, ingress, identity, monitoring, developer platform tooling and application delivery. The role will play a key part in improving the operational maturity of our platform estate, supporting the migration from VMware to KVM, strengthening GitOps and CI/CD practices, and helping ensure our platforms remain secure, scalable and aligned to the needs of regulated customer environments. You'll work closely with platform, application, AI, networking, security, QA and architecture teams to build reliable foundations that enable other engineering teams to deliver safely and at pace. This is not a ticket handling role. It is a senior engineering position where you'll be expected to own problems, drive improvements, and help shape how TIG operates critical cloud native infrastructure. About the Team You'll be joining our Cloud team, working closely with Platform Engineering and wider engineering teams responsible for the foundational platforms on which TIG's services run. This is a great opportunity to join a small, senior technical environment where you can have direct ownership, meaningful influence, and visibility across modern platform engineering, Kubernetes, automation, observability, security and cloud native delivery. Key Responsibilities Operate, harden and extend production OpenShift / OKD / Kubernetes clusters across on premises and hybrid environments. Support the migration from VMware to KVM, helping modernise the underlying compute and storage layer. Own and improve CI/CD processes across the full lifecycle of platform and application components. Work with platform and application engineers to support cloud native delivery using tools such as Helm and Kustomize. Develop and mature GitOps deployment practices using tools such as Argo CD or Flux. Maintain and improve core platform services including identity, ingress, observability, certificate management, service mesh and container registry capabilities. Build and operate observability across logs, metrics, traces, alerting, SLOs and error budgets. Improve platform hardening in line with secure and regulated environment requirements, including network policy, SELinux, image provenance, secret management and audit. Automate repeatable operational tasks using tools such as Ansible, Terraform, Helm, Kustomize, Go, Python or equivalent technologies. Lead incident response activity, support blameless post mortems and drive systemic fixes. Partner with networking and security teams on platform integration, segmentation, load balancing and accreditation evidence. Create and maintain clear technical documentation, runbooks, design notes and operational guidance. Mentor other engineers and act as a senior technical authority across cloud and Kubernetes operations. Participate in an on call rota, with appropriate compensation. Success in This Role Looks Like A more reliable, secure and measurable production Kubernetes estate. Improved platform observability, with meaningful alerting, SLOs and trend data that engineering teams actively use. Progress against the VMware to KVM migration, with a clear and automated path for the underlying infrastructure layer. A mature GitOps approach covering platform and application components, including rollback, drift detection and operational control. Improved CI/CD practices that help teams move at pace while considering security, QA and compliance earlier in the lifecycle. Well documented, supportable and scalable platform services. Stronger incident response, clearer runbooks and post mortems that lead to real operational improvements. Recognition as a technical authority for Kubernetes, cloud and platform operations across the organisation. What We're Looking For We're looking for a Senior Site Reliability Engineer (SRE) with strong experience operating production Kubernetes environments. This role is well suited to someone who combines deep technical capability with strong operational discipline. You'll be comfortable taking ownership of complex platform challenges, improving reliability, and working collaboratively across engineering, security, networking and architecture teams. Essential Experience & Skills Strong experience running production Kubernetes environments, not just consuming or deploying into them. Strong Linux fundamentals, including systemd, networking, storage and performance troubleshooting. Experience with at least one Kubernetes distribution such as OKD, OpenShift, vanilla Kubernetes, Rancher, EKS, AKS or GKE. Solid infrastructure as code experience, including Ansible plus Terraform or equivalent, alongside tools such as Helm and Kustomize. GitOps and CI/CD experience managing full application and component lifecycles, using tools such as Argo CD, Flux, GitHub Actions or similar. Prometheus, Grafana, Elastic Stack / LGTM, OpenTelemetry or similar. Experience working with identity and access technologies such as OIDC, SAML, SCIM or Keycloak. Experience with virtualisation or infrastructure platforms such as KVM, libvirt or VMware. Scripting or tooling experience using Go, Python, shell scripting or similar. Strong troubleshooting, problem solving and analytical skills. Experience working in secure, regulated or enterprise scale environments. Strong communication skills, with the ability to produce clear documentation, runbooks, post mortems and technical guidance. Eligible to hold UK SC clearance. Desirable (Not Essential) Specific OpenShift or OKD experience, including operators, MachineConfig or SCCs. Service mesh experience such as Istio or Linkerd. Policy engine experience such as OPA, Gatekeeper or Kyverno. Cloud native application deployment experience using Helm, Terraform, Kustomize or similar. Storage experience such as Ceph, Longhorn, OpenShift Data Foundation or equivalent. Networking experience including BGP, VXLAN, Palo Alto or Juniper technologies. Software supply chain security experience, including SBOMs, image signing, admission control or tools such as Sigstore. Experience operating AI, ML or GPU enabled platforms. CKA, CKAD, CKS, Red Hat certifications or equivalent. Active or recent UK SC clearance. Recognised open source contributions to the Kubernetes ecosystem. Soft Skills & Behaviours Calm, structured and methodical under pressure. Strong written and verbal communication skills. Collaborative working style across platform, development, QA, security, networking and architecture teams. Strong sense of ownership and accountability. Automation first mindset, with a focus on removing repeatable manual work. Able to influence technical practice through evidence, example and credibility. Pragmatic and solutions focused approach to problem solving. Curious about why systems fail, not just how to bring them back online. . click apply for full job details
25/07/2026
Full time
Role: Senior Site Reliability Engineer (SRE) - Kubernetes / OpenShift Location: Remote - UK (possible paid occasional travel to TIG Secure site locations as required) Job Type: Full-time, Permanent (37.5 hours) Salary: Competitive + benefits + package Security Clearance Requirements Please note that holding a current Security Clearance is not essential at the time of application, but eligibility is required. This role requires the successful candidate to be eligible for Security Check (SC) clearance. To meet this requirement, applicants must: Have the right to work in the UK Have lived in the UK continuously for the past 5 years Not have spent more than 6 months outside the UK in total during that period Be willing to undergo security vetting as part of the onboarding process About You You're an experienced SRE, Platform Engineer or Cloud Engineer with strong hands on experience running Kubernetes in production environments. You're comfortable working across Linux, Kubernetes, cloud native tooling, automation, observability, CI/CD and infrastructure as code. You understand that reliability, security and operational maturity are critical to how modern platforms support engineering teams and customer facing services. You enjoy treating infrastructure as a product, automating repeatable work, improving resilience, and building platforms that other engineers can rely on. You're calm under pressure, methodical during incidents, and able to turn operational challenges into long term improvements. You may have worked in a regulated, secure, government, defence, financial services, telecoms, managed services or cloud native environment, but most importantly you have operated Kubernetes at depth and understand the realities of production ownership. You're a senior individual contributor who can mentor others, influence engineering practice, and provide technical authority without needing formal line management responsibility. About the Role We're looking for a Senior Site Reliability Engineer (SRE) to help operate, harden and mature our production OKD / Kubernetes platforms. This is a hands on engineering role focused on reliability, automation, observability, GitOps, CI/CD and secure platform operations. You'll work across the full stack, from bare metal and virtualisation through to Kubernetes control plane operations, ingress, identity, monitoring, developer platform tooling and application delivery. The role will play a key part in improving the operational maturity of our platform estate, supporting the migration from VMware to KVM, strengthening GitOps and CI/CD practices, and helping ensure our platforms remain secure, scalable and aligned to the needs of regulated customer environments. You'll work closely with platform, application, AI, networking, security, QA and architecture teams to build reliable foundations that enable other engineering teams to deliver safely and at pace. This is not a ticket handling role. It is a senior engineering position where you'll be expected to own problems, drive improvements, and help shape how TIG operates critical cloud native infrastructure. About the Team You'll be joining our Cloud team, working closely with Platform Engineering and wider engineering teams responsible for the foundational platforms on which TIG's services run. This is a great opportunity to join a small, senior technical environment where you can have direct ownership, meaningful influence, and visibility across modern platform engineering, Kubernetes, automation, observability, security and cloud native delivery. Key Responsibilities Operate, harden and extend production OpenShift / OKD / Kubernetes clusters across on premises and hybrid environments. Support the migration from VMware to KVM, helping modernise the underlying compute and storage layer. Own and improve CI/CD processes across the full lifecycle of platform and application components. Work with platform and application engineers to support cloud native delivery using tools such as Helm and Kustomize. Develop and mature GitOps deployment practices using tools such as Argo CD or Flux. Maintain and improve core platform services including identity, ingress, observability, certificate management, service mesh and container registry capabilities. Build and operate observability across logs, metrics, traces, alerting, SLOs and error budgets. Improve platform hardening in line with secure and regulated environment requirements, including network policy, SELinux, image provenance, secret management and audit. Automate repeatable operational tasks using tools such as Ansible, Terraform, Helm, Kustomize, Go, Python or equivalent technologies. Lead incident response activity, support blameless post mortems and drive systemic fixes. Partner with networking and security teams on platform integration, segmentation, load balancing and accreditation evidence. Create and maintain clear technical documentation, runbooks, design notes and operational guidance. Mentor other engineers and act as a senior technical authority across cloud and Kubernetes operations. Participate in an on call rota, with appropriate compensation. Success in This Role Looks Like A more reliable, secure and measurable production Kubernetes estate. Improved platform observability, with meaningful alerting, SLOs and trend data that engineering teams actively use. Progress against the VMware to KVM migration, with a clear and automated path for the underlying infrastructure layer. A mature GitOps approach covering platform and application components, including rollback, drift detection and operational control. Improved CI/CD practices that help teams move at pace while considering security, QA and compliance earlier in the lifecycle. Well documented, supportable and scalable platform services. Stronger incident response, clearer runbooks and post mortems that lead to real operational improvements. Recognition as a technical authority for Kubernetes, cloud and platform operations across the organisation. What We're Looking For We're looking for a Senior Site Reliability Engineer (SRE) with strong experience operating production Kubernetes environments. This role is well suited to someone who combines deep technical capability with strong operational discipline. You'll be comfortable taking ownership of complex platform challenges, improving reliability, and working collaboratively across engineering, security, networking and architecture teams. Essential Experience & Skills Strong experience running production Kubernetes environments, not just consuming or deploying into them. Strong Linux fundamentals, including systemd, networking, storage and performance troubleshooting. Experience with at least one Kubernetes distribution such as OKD, OpenShift, vanilla Kubernetes, Rancher, EKS, AKS or GKE. Solid infrastructure as code experience, including Ansible plus Terraform or equivalent, alongside tools such as Helm and Kustomize. GitOps and CI/CD experience managing full application and component lifecycles, using tools such as Argo CD, Flux, GitHub Actions or similar. Prometheus, Grafana, Elastic Stack / LGTM, OpenTelemetry or similar. Experience working with identity and access technologies such as OIDC, SAML, SCIM or Keycloak. Experience with virtualisation or infrastructure platforms such as KVM, libvirt or VMware. Scripting or tooling experience using Go, Python, shell scripting or similar. Strong troubleshooting, problem solving and analytical skills. Experience working in secure, regulated or enterprise scale environments. Strong communication skills, with the ability to produce clear documentation, runbooks, post mortems and technical guidance. Eligible to hold UK SC clearance. Desirable (Not Essential) Specific OpenShift or OKD experience, including operators, MachineConfig or SCCs. Service mesh experience such as Istio or Linkerd. Policy engine experience such as OPA, Gatekeeper or Kyverno. Cloud native application deployment experience using Helm, Terraform, Kustomize or similar. Storage experience such as Ceph, Longhorn, OpenShift Data Foundation or equivalent. Networking experience including BGP, VXLAN, Palo Alto or Juniper technologies. Software supply chain security experience, including SBOMs, image signing, admission control or tools such as Sigstore. Experience operating AI, ML or GPU enabled platforms. CKA, CKAD, CKS, Red Hat certifications or equivalent. Active or recent UK SC clearance. Recognised open source contributions to the Kubernetes ecosystem. Soft Skills & Behaviours Calm, structured and methodical under pressure. Strong written and verbal communication skills. Collaborative working style across platform, development, QA, security, networking and architecture teams. Strong sense of ownership and accountability. Automation first mindset, with a focus on removing repeatable manual work. Able to influence technical practice through evidence, example and credibility. Pragmatic and solutions focused approach to problem solving. Curious about why systems fail, not just how to bring them back online. . click apply for full job details
This is not a live vacancy. We are building a talent pipeline ahead of anticipated demand and are keen to connect with strong Senior and Lead Cloud Platform Engineers for future opportunities. Role Overview We are seeking an experienced Cloud Platform Engineer to join our platform engineering team. This is a hands on role focused on designing, building, operating, and continuously improving cloud native platforms that enable development teams to deliver reliable, scalable, and secure applications. The successful candidate will combine deep expertise in Kubernetes, cloud infrastructure, automation, and DevOps practices with a strong operational mindset. They will be responsible for maintaining the health and evolution of containerised workloads, enhancing developer experience, and driving improvements across infrastructure, deployment pipelines, and platform services. At the Lead level, this role also combines hands on expertise with team leadership, stakeholder engagement, and delivery planning. The right candidate can operate credibly at an engineering level while also communicating clearly with business and technical stakeholders, translating platform priorities into a coherent roadmap and ensuring the team has the clarity and support needed to deliver effectively. This role requires a proactive engineer who can work independently, collaborate effectively across teams, and contribute to the ongoing maturity of modern cloud platforms. Key Responsibilities Own and support the operational health, reliability, scalability, and performance of Kubernetes based platforms and containerised workloads Manage and evolve Kubernetes environments, including workload deployment, configuration management, cluster operations, and platform upgrades Drive and maintain GitOps practices to ensure consistent, automated, and auditable infrastructure and application deployments Build, maintain, and improve infrastructure as code solutions to support scalable, repeatable, and secure cloud environments Develop and enhance CI/CD pipelines to streamline software delivery, improve deployment reliability, and reduce operational overhead Manage and optimise cloud native and managed platform services, including capacity planning, performance tuning, cost optimisation, patching, and lifecycle management Partner with software engineering teams to support application deployment, troubleshooting, observability, and platform adoption Monitor platform health and respond to incidents, conducting root cause analysis and implementing preventative improvements Identify technical debt and contribute to platform roadmaps focused on operational excellence, security, automation, and developer productivity Champion best practices in cloud architecture, security, reliability, automation, and platform engineering Participate in on call support and incident management activities as required, including for the EKS platform Additional Key Responsibilities for the Lead Own the platform engineering roadmap, working with business and technical stakeholders to prioritise initiatives, surface risks, and align delivery against wider programme goals Lead and mentor a team of platform engineers, providing technical guidance, supporting growth, and maintaining quality and consistency across the team's output Act as the primary technical point of contact for platform related matters across engineering, product, and operations teams Translate platform priorities into clear, actionable plans, ensuring the team understands objectives, dependencies, and delivery timelines Contribute to resourcing, planning, and prioritisation discussions, balancing technical debt, reliability, and new feature delivery Provide regular updates on roadmap progress, risks, and dependencies to relevant stakeholders and leadership Skills & Experience Strong experience operating Kubernetes in production environments, including cluster administration, workload management, networking, security, and performance optimisation Experience with cloud platforms such as AWS, Azure, or Google Cloud Platform, with a strong understanding of managed cloud services Hands on experience with GitOps methodologies and related tooling Experience with infrastructure as code tools such as Terraform Strong understanding of CI/CD principles and experience building and maintaining automated delivery pipelines Experience working with container technologies and cloud native architectures Knowledge of observability, monitoring, logging, and incident management practices Strong troubleshooting and problem solving skills across infrastructure, platform, and application layers Experience supporting software development teams in deploying and operating distributed applications Ability to work independently while collaborating effectively with cross functional teams Betting and gaming domain knowledge required Additional Essential Skills for the Lead Demonstrated experience leading or mentoring engineers, with the ability to balance technical leadership with people development Strong stakeholder management skills, with the ability to translate business priorities into technical roadmaps and vice versa Excellent communication skills, comfortable operating across both technical and non technical audiences Experience with Amazon EKS or other managed Kubernetes platforms Familiarity with Flux, Argo CD, Helm, and Kustomize Experience supporting managed data and messaging platforms such as RDS, Kafka, or equivalent services Knowledge of application runtimes and services built using Java, Kotlin, Python, or similar technologies Experience implementing cloud security controls, governance, and compliance requirements Exposure to site reliability engineering (SRE) practices and platform engineering frameworks What we'll offer you: We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You'll be part of a supportive, and friendly culture, surrounded by smart, curious people who care deeply about what they do. We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed. We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.
25/07/2026
Full time
This is not a live vacancy. We are building a talent pipeline ahead of anticipated demand and are keen to connect with strong Senior and Lead Cloud Platform Engineers for future opportunities. Role Overview We are seeking an experienced Cloud Platform Engineer to join our platform engineering team. This is a hands on role focused on designing, building, operating, and continuously improving cloud native platforms that enable development teams to deliver reliable, scalable, and secure applications. The successful candidate will combine deep expertise in Kubernetes, cloud infrastructure, automation, and DevOps practices with a strong operational mindset. They will be responsible for maintaining the health and evolution of containerised workloads, enhancing developer experience, and driving improvements across infrastructure, deployment pipelines, and platform services. At the Lead level, this role also combines hands on expertise with team leadership, stakeholder engagement, and delivery planning. The right candidate can operate credibly at an engineering level while also communicating clearly with business and technical stakeholders, translating platform priorities into a coherent roadmap and ensuring the team has the clarity and support needed to deliver effectively. This role requires a proactive engineer who can work independently, collaborate effectively across teams, and contribute to the ongoing maturity of modern cloud platforms. Key Responsibilities Own and support the operational health, reliability, scalability, and performance of Kubernetes based platforms and containerised workloads Manage and evolve Kubernetes environments, including workload deployment, configuration management, cluster operations, and platform upgrades Drive and maintain GitOps practices to ensure consistent, automated, and auditable infrastructure and application deployments Build, maintain, and improve infrastructure as code solutions to support scalable, repeatable, and secure cloud environments Develop and enhance CI/CD pipelines to streamline software delivery, improve deployment reliability, and reduce operational overhead Manage and optimise cloud native and managed platform services, including capacity planning, performance tuning, cost optimisation, patching, and lifecycle management Partner with software engineering teams to support application deployment, troubleshooting, observability, and platform adoption Monitor platform health and respond to incidents, conducting root cause analysis and implementing preventative improvements Identify technical debt and contribute to platform roadmaps focused on operational excellence, security, automation, and developer productivity Champion best practices in cloud architecture, security, reliability, automation, and platform engineering Participate in on call support and incident management activities as required, including for the EKS platform Additional Key Responsibilities for the Lead Own the platform engineering roadmap, working with business and technical stakeholders to prioritise initiatives, surface risks, and align delivery against wider programme goals Lead and mentor a team of platform engineers, providing technical guidance, supporting growth, and maintaining quality and consistency across the team's output Act as the primary technical point of contact for platform related matters across engineering, product, and operations teams Translate platform priorities into clear, actionable plans, ensuring the team understands objectives, dependencies, and delivery timelines Contribute to resourcing, planning, and prioritisation discussions, balancing technical debt, reliability, and new feature delivery Provide regular updates on roadmap progress, risks, and dependencies to relevant stakeholders and leadership Skills & Experience Strong experience operating Kubernetes in production environments, including cluster administration, workload management, networking, security, and performance optimisation Experience with cloud platforms such as AWS, Azure, or Google Cloud Platform, with a strong understanding of managed cloud services Hands on experience with GitOps methodologies and related tooling Experience with infrastructure as code tools such as Terraform Strong understanding of CI/CD principles and experience building and maintaining automated delivery pipelines Experience working with container technologies and cloud native architectures Knowledge of observability, monitoring, logging, and incident management practices Strong troubleshooting and problem solving skills across infrastructure, platform, and application layers Experience supporting software development teams in deploying and operating distributed applications Ability to work independently while collaborating effectively with cross functional teams Betting and gaming domain knowledge required Additional Essential Skills for the Lead Demonstrated experience leading or mentoring engineers, with the ability to balance technical leadership with people development Strong stakeholder management skills, with the ability to translate business priorities into technical roadmaps and vice versa Excellent communication skills, comfortable operating across both technical and non technical audiences Experience with Amazon EKS or other managed Kubernetes platforms Familiarity with Flux, Argo CD, Helm, and Kustomize Experience supporting managed data and messaging platforms such as RDS, Kafka, or equivalent services Knowledge of application runtimes and services built using Java, Kotlin, Python, or similar technologies Experience implementing cloud security controls, governance, and compliance requirements Exposure to site reliability engineering (SRE) practices and platform engineering frameworks What we'll offer you: We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You'll be part of a supportive, and friendly culture, surrounded by smart, curious people who care deeply about what they do. We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed. We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.
This is not a live vacancy. We are building a talent pipeline ahead of anticipated demand and are keen to connect with strong Senior and Lead Cloud Platform Engineers for future opportunities. Role Overview We are seeking an experienced Cloud Platform Engineer to join our platform engineering team. This is a hands on role focused on designing, building, operating, and continuously improving cloud native platforms that enable development teams to deliver reliable, scalable, and secure applications. The successful candidate will combine deep expertise in Kubernetes, cloud infrastructure, automation, and DevOps practices with a strong operational mindset. They will be responsible for maintaining the health and evolution of containerised workloads, enhancing developer experience, and driving improvements across infrastructure, deployment pipelines, and platform services. At the Lead level, this role also combines hands on expertise with team leadership, stakeholder engagement, and delivery planning. The right candidate can operate credibly at an engineering level while also communicating clearly with business and technical stakeholders, translating platform priorities into a coherent roadmap and ensuring the team has the clarity and support needed to deliver effectively. This role requires a proactive engineer who can work independently, collaborate effectively across teams, and contribute to the ongoing maturity of modern cloud platforms. Key Responsibilities Own and support the operational health, reliability, scalability, and performance of Kubernetes based platforms and containerised workloads Manage and evolve Kubernetes environments, including workload deployment, configuration management, cluster operations, and platform upgrades Drive and maintain GitOps practices to ensure consistent, automated, and auditable infrastructure and application deployments Build, maintain, and improve infrastructure as code solutions to support scalable, repeatable, and secure cloud environments Develop and enhance CI/CD pipelines to streamline software delivery, improve deployment reliability, and reduce operational overhead Manage and optimise cloud native and managed platform services, including capacity planning, performance tuning, cost optimisation, patching, and lifecycle management Partner with software engineering teams to support application deployment, troubleshooting, observability, and platform adoption Monitor platform health and respond to incidents, conducting root cause analysis and implementing preventative improvements Identify technical debt and contribute to platform roadmaps focused on operational excellence, security, automation, and developer productivity Champion best practices in cloud architecture, security, reliability, automation, and platform engineering Participate in on call support and incident management activities as required, including for the EKS platform Additional Key Responsibilities for the Lead Own the platform engineering roadmap, working with business and technical stakeholders to prioritise initiatives, surface risks, and align delivery against wider programme goals Lead and mentor a team of platform engineers, providing technical guidance, supporting growth, and maintaining quality and consistency across the team's output Act as the primary technical point of contact for platform related matters across engineering, product, and operations teams Translate platform priorities into clear, actionable plans, ensuring the team understands objectives, dependencies, and delivery timelines Contribute to resourcing, planning, and prioritisation discussions, balancing technical debt, reliability, and new feature delivery Provide regular updates on roadmap progress, risks, and dependencies to relevant stakeholders and leadership Skills & Experience Strong experience operating Kubernetes in production environments, including cluster administration, workload management, networking, security, and performance optimisation Experience with cloud platforms such as AWS, Azure, or Google Cloud Platform, with a strong understanding of managed cloud services Hands on experience with GitOps methodologies and related tooling Experience with infrastructure as code tools such as Terraform Strong understanding of CI/CD principles and experience building and maintaining automated delivery pipelines Experience working with container technologies and cloud native architectures Knowledge of observability, monitoring, logging, and incident management practices Strong troubleshooting and problem solving skills across infrastructure, platform, and application layers Experience supporting software development teams in deploying and operating distributed applications Ability to work independently while collaborating effectively with cross functional teams Betting and gaming domain knowledge required Additional Essential Skills for the Lead Demonstrated experience leading or mentoring engineers, with the ability to balance technical leadership with people development Strong stakeholder management skills, with the ability to translate business priorities into technical roadmaps and vice versa Excellent communication skills, comfortable operating across both technical and non technical audiences Experience with Amazon EKS or other managed Kubernetes platforms Familiarity with Flux, Argo CD, Helm, and Kustomize Experience supporting managed data and messaging platforms such as RDS, Kafka, or equivalent services Knowledge of application runtimes and services built using Java, Kotlin, Python, or similar technologies Experience implementing cloud security controls, governance, and compliance requirements Exposure to site reliability engineering (SRE) practices and platform engineering frameworks What we'll offer you: We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You'll be part of a supportive, and friendly culture, surrounded by smart, curious people who care deeply about what they do. We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed. We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.
25/07/2026
Full time
This is not a live vacancy. We are building a talent pipeline ahead of anticipated demand and are keen to connect with strong Senior and Lead Cloud Platform Engineers for future opportunities. Role Overview We are seeking an experienced Cloud Platform Engineer to join our platform engineering team. This is a hands on role focused on designing, building, operating, and continuously improving cloud native platforms that enable development teams to deliver reliable, scalable, and secure applications. The successful candidate will combine deep expertise in Kubernetes, cloud infrastructure, automation, and DevOps practices with a strong operational mindset. They will be responsible for maintaining the health and evolution of containerised workloads, enhancing developer experience, and driving improvements across infrastructure, deployment pipelines, and platform services. At the Lead level, this role also combines hands on expertise with team leadership, stakeholder engagement, and delivery planning. The right candidate can operate credibly at an engineering level while also communicating clearly with business and technical stakeholders, translating platform priorities into a coherent roadmap and ensuring the team has the clarity and support needed to deliver effectively. This role requires a proactive engineer who can work independently, collaborate effectively across teams, and contribute to the ongoing maturity of modern cloud platforms. Key Responsibilities Own and support the operational health, reliability, scalability, and performance of Kubernetes based platforms and containerised workloads Manage and evolve Kubernetes environments, including workload deployment, configuration management, cluster operations, and platform upgrades Drive and maintain GitOps practices to ensure consistent, automated, and auditable infrastructure and application deployments Build, maintain, and improve infrastructure as code solutions to support scalable, repeatable, and secure cloud environments Develop and enhance CI/CD pipelines to streamline software delivery, improve deployment reliability, and reduce operational overhead Manage and optimise cloud native and managed platform services, including capacity planning, performance tuning, cost optimisation, patching, and lifecycle management Partner with software engineering teams to support application deployment, troubleshooting, observability, and platform adoption Monitor platform health and respond to incidents, conducting root cause analysis and implementing preventative improvements Identify technical debt and contribute to platform roadmaps focused on operational excellence, security, automation, and developer productivity Champion best practices in cloud architecture, security, reliability, automation, and platform engineering Participate in on call support and incident management activities as required, including for the EKS platform Additional Key Responsibilities for the Lead Own the platform engineering roadmap, working with business and technical stakeholders to prioritise initiatives, surface risks, and align delivery against wider programme goals Lead and mentor a team of platform engineers, providing technical guidance, supporting growth, and maintaining quality and consistency across the team's output Act as the primary technical point of contact for platform related matters across engineering, product, and operations teams Translate platform priorities into clear, actionable plans, ensuring the team understands objectives, dependencies, and delivery timelines Contribute to resourcing, planning, and prioritisation discussions, balancing technical debt, reliability, and new feature delivery Provide regular updates on roadmap progress, risks, and dependencies to relevant stakeholders and leadership Skills & Experience Strong experience operating Kubernetes in production environments, including cluster administration, workload management, networking, security, and performance optimisation Experience with cloud platforms such as AWS, Azure, or Google Cloud Platform, with a strong understanding of managed cloud services Hands on experience with GitOps methodologies and related tooling Experience with infrastructure as code tools such as Terraform Strong understanding of CI/CD principles and experience building and maintaining automated delivery pipelines Experience working with container technologies and cloud native architectures Knowledge of observability, monitoring, logging, and incident management practices Strong troubleshooting and problem solving skills across infrastructure, platform, and application layers Experience supporting software development teams in deploying and operating distributed applications Ability to work independently while collaborating effectively with cross functional teams Betting and gaming domain knowledge required Additional Essential Skills for the Lead Demonstrated experience leading or mentoring engineers, with the ability to balance technical leadership with people development Strong stakeholder management skills, with the ability to translate business priorities into technical roadmaps and vice versa Excellent communication skills, comfortable operating across both technical and non technical audiences Experience with Amazon EKS or other managed Kubernetes platforms Familiarity with Flux, Argo CD, Helm, and Kustomize Experience supporting managed data and messaging platforms such as RDS, Kafka, or equivalent services Knowledge of application runtimes and services built using Java, Kotlin, Python, or similar technologies Experience implementing cloud security controls, governance, and compliance requirements Exposure to site reliability engineering (SRE) practices and platform engineering frameworks What we'll offer you: We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You'll be part of a supportive, and friendly culture, surrounded by smart, curious people who care deeply about what they do. We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed. We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.
Cloud Platform Engineer (Senior / Lead) Department: Technology Employment Type: Permanent - Full Time Location: London Reporting To: Description About the role We are looking for a hands on Senior / Lead Cloud Platform Engineer to shape and run the multi cloud foundation that our engineering, data and product teams build on everyday. This is a senior individual contributor / tech lead role at the heart of a modern hybrid cloud strategy. Our goal is to make cloud infrastructure effectively invisible and commoditised for internal teams: engineers should ship through self service golden paths and a well designed internal developer platform, without needing to understand the underlying cloud plumbing. You will treat platform capabilities as products, with clear APIs, paved roads, strong defaults and great developer experience. A defining part of this role is preparing our infrastructure for a future (or present) where some "engineers" are AI agents. You will design platforms, guardrails and interfaces that let both human engineers and autonomous AI agents provision, operate and optimise infrastructure safely - and you will use AI agents heavily yourself to automate and optimise day to day platform work. Security is central, not an afterthought. You will help drive a zero trust approach across identity, network, workloads and data, ensuring the platform is secure by default for humans and machine/agent identities alike. You will also work closely with our data function, helping design and optimise the data pipelines, data stores and large scale analytics infrastructure (for example BigQuery and similar warehouses) that the business depends on. What you'll do Design, build and operate our multi cloud and hybrid cloud platform across at least two of the top three providers (AWS, Azure and/or Google Cloud), plus on prem/hybrid connectivity where needed. Build and own an internal developer platform and self service "golden paths" that make cloud infrastructure feel invisible and commoditised for engineering, data and product teams; and their AI agents. Deliver everything as code: infrastructure as code, GitOps, reusable modules, CI/CD pipelines and policy as code guardrails. Leverage AI agents extensively to automate and optimise platform work-provisioning, cost and performance optimisation, incident response, remediation and documentation. Prepare the infrastructure for AI agents as first class "engineers": safe machine identities, scoped permissions, sandboxes, approval workflows and audit trails so agents can provision and operate infrastructure within tight guardrails. Embed a zero trust security model across identity, network, workloads and data for both human and machine/agent identities; secure by default, least privilege, secrets management and continuous compliance. Apply SRE practices-SLOs/SLIs, observability, capacity planning, resilience and blameless incident management-to keep the platform reliable and cost efficient. Partner with data engineering to design and optimise data pipelines, data stores and large scale analytics infrastructure such as BigQuery, including query, cost and performance tuning. Mentor engineers, set technical direction and champion strong platform and security engineering standards across the organisation. What you'll need to succeed Essential requirements: Extensive hands on experience designing, building and operating production cloud infrastructure at senior or lead level. Multi cloud experience across at least two of the top three providers (AWS, Microsoft Azure and Google Cloud), including a recognised professional level cloud certification for each of those two providers (for example AWS Solutions Architect / DevOps Engineer Professional, Azure Solutions Architect / DevOps Engineer Expert, or Google Cloud Professional Cloud Architect / DevOps Engineer). Strong background in modern hybrid cloud architecture and connecting cloud with on prem/edge environments. Deep infrastructure as code and automation skills (e.g. Terraform/OpenTofu, Pulumi, Ansible), GitOps and CI/CD, plus containers and orchestration (Docker, Kubernetes). Proven experience building internal developer platforms, self service golden paths and platform as a product to abstract away cloud complexity for engineering teams. Practical experience using AI agents / LLM based tooling to automate and optimise infrastructure work, and interest in designing infrastructure that AI agents can operate safely. Strong security engineering mindset with hands on zero trust experience across identity, network, workloads and data-including secrets management, least privilege IAM and machine/workload identity. Solid programming/scripting ability (e.g. Python, Go) and strong observability, reliability and cost optimisation practices. Desirable requirements: Experience working as a Site Reliability Engineer (SRE) with SLOs/SLIs, error budgets and incident management. A third top tier cloud certification, or specialist security/Kubernetes certifications (e.g. CKA/CKS). Significant data engineering experience: designing and operating data pipelines and data stores, and optimising databases and large scale data infrastructure such as BigQuery (including query, cost and performance tuning). Experience preparing environments for autonomous or agentic workloads-sandboxes, scoped machine identities, approval workflows and audit trails. Experience in a regulated or fintech environment. Your approach to work: Pragmatic and hands on, with a strong bias for automation and eliminating toil. Product mindset- you treat internal engineers (human and AI) as your customers and obsess over their experience. Security and reliability first, collaborative, and comfortable leading and mentoring. Important to know Location: We have multiple offices across the UK. We have a new office in London which is becoming more central to where we collaborate in person. We have a flexible working policy with a few days per week in the office. Right to Work: Applicants must already hold a legal right to work in the UK without time restrictions and without the need for future sponsorship. We are unable to provide Skilled Worker visa sponsorship.
25/07/2026
Full time
Cloud Platform Engineer (Senior / Lead) Department: Technology Employment Type: Permanent - Full Time Location: London Reporting To: Description About the role We are looking for a hands on Senior / Lead Cloud Platform Engineer to shape and run the multi cloud foundation that our engineering, data and product teams build on everyday. This is a senior individual contributor / tech lead role at the heart of a modern hybrid cloud strategy. Our goal is to make cloud infrastructure effectively invisible and commoditised for internal teams: engineers should ship through self service golden paths and a well designed internal developer platform, without needing to understand the underlying cloud plumbing. You will treat platform capabilities as products, with clear APIs, paved roads, strong defaults and great developer experience. A defining part of this role is preparing our infrastructure for a future (or present) where some "engineers" are AI agents. You will design platforms, guardrails and interfaces that let both human engineers and autonomous AI agents provision, operate and optimise infrastructure safely - and you will use AI agents heavily yourself to automate and optimise day to day platform work. Security is central, not an afterthought. You will help drive a zero trust approach across identity, network, workloads and data, ensuring the platform is secure by default for humans and machine/agent identities alike. You will also work closely with our data function, helping design and optimise the data pipelines, data stores and large scale analytics infrastructure (for example BigQuery and similar warehouses) that the business depends on. What you'll do Design, build and operate our multi cloud and hybrid cloud platform across at least two of the top three providers (AWS, Azure and/or Google Cloud), plus on prem/hybrid connectivity where needed. Build and own an internal developer platform and self service "golden paths" that make cloud infrastructure feel invisible and commoditised for engineering, data and product teams; and their AI agents. Deliver everything as code: infrastructure as code, GitOps, reusable modules, CI/CD pipelines and policy as code guardrails. Leverage AI agents extensively to automate and optimise platform work-provisioning, cost and performance optimisation, incident response, remediation and documentation. Prepare the infrastructure for AI agents as first class "engineers": safe machine identities, scoped permissions, sandboxes, approval workflows and audit trails so agents can provision and operate infrastructure within tight guardrails. Embed a zero trust security model across identity, network, workloads and data for both human and machine/agent identities; secure by default, least privilege, secrets management and continuous compliance. Apply SRE practices-SLOs/SLIs, observability, capacity planning, resilience and blameless incident management-to keep the platform reliable and cost efficient. Partner with data engineering to design and optimise data pipelines, data stores and large scale analytics infrastructure such as BigQuery, including query, cost and performance tuning. Mentor engineers, set technical direction and champion strong platform and security engineering standards across the organisation. What you'll need to succeed Essential requirements: Extensive hands on experience designing, building and operating production cloud infrastructure at senior or lead level. Multi cloud experience across at least two of the top three providers (AWS, Microsoft Azure and Google Cloud), including a recognised professional level cloud certification for each of those two providers (for example AWS Solutions Architect / DevOps Engineer Professional, Azure Solutions Architect / DevOps Engineer Expert, or Google Cloud Professional Cloud Architect / DevOps Engineer). Strong background in modern hybrid cloud architecture and connecting cloud with on prem/edge environments. Deep infrastructure as code and automation skills (e.g. Terraform/OpenTofu, Pulumi, Ansible), GitOps and CI/CD, plus containers and orchestration (Docker, Kubernetes). Proven experience building internal developer platforms, self service golden paths and platform as a product to abstract away cloud complexity for engineering teams. Practical experience using AI agents / LLM based tooling to automate and optimise infrastructure work, and interest in designing infrastructure that AI agents can operate safely. Strong security engineering mindset with hands on zero trust experience across identity, network, workloads and data-including secrets management, least privilege IAM and machine/workload identity. Solid programming/scripting ability (e.g. Python, Go) and strong observability, reliability and cost optimisation practices. Desirable requirements: Experience working as a Site Reliability Engineer (SRE) with SLOs/SLIs, error budgets and incident management. A third top tier cloud certification, or specialist security/Kubernetes certifications (e.g. CKA/CKS). Significant data engineering experience: designing and operating data pipelines and data stores, and optimising databases and large scale data infrastructure such as BigQuery (including query, cost and performance tuning). Experience preparing environments for autonomous or agentic workloads-sandboxes, scoped machine identities, approval workflows and audit trails. Experience in a regulated or fintech environment. Your approach to work: Pragmatic and hands on, with a strong bias for automation and eliminating toil. Product mindset- you treat internal engineers (human and AI) as your customers and obsess over their experience. Security and reliability first, collaborative, and comfortable leading and mentoring. Important to know Location: We have multiple offices across the UK. We have a new office in London which is becoming more central to where we collaborate in person. We have a flexible working policy with a few days per week in the office. Right to Work: Applicants must already hold a legal right to work in the UK without time restrictions and without the need for future sponsorship. We are unable to provide Skilled Worker visa sponsorship.
GCP Platform Engineer Location: Hybrid / Manchester Salary: Up to £80,000 Benefits Type: Permanent We're partnering with an innovative platform business that is investing heavily in its cloud infrastructure and engineering capabilities. They are looking for an experienced GCP Platform Engineer to help build, automate and scale a modern cloud platform used across multiple products and teams. This is an opportunity to join a collaborative engineering environment where infrastructure is treated as code, automation is a priority, and engineers are empowered to drive technical decisions. The Role As a GCP Platform Engineer, you will be responsible for designing, building and maintaining the company's cloud platform on Google Cloud Platform (GCP). You'll work closely with software engineers and DevOps teams to create scalable, secure and highly available infrastructure. Key responsibilities Designing and implementing cloud infrastructure on GCP Building and maintaining Infrastructure as Code using Terraform Automating infrastructure provisioning and deployment pipelines Managing Kubernetes and containerised workloads Implementing monitoring, logging and observability solutions Driving platform reliability, security and best practices Collaborating with engineering teams to improve developer experience Skills & Experience Essential: Strong commercial experience with Google Cloud Platform (GCP) Extensive experience with Terraform and Infrastructure as Code Experience building CI/CD pipelines Knowledge of Kubernetes / GKE and container technologies Experience with Linux and scripting (Bash, Python or Go) Understanding of networking, IAM and cloud security principles Experience with monitoring and observability tooling Desirable: Experience with GitOps practices Knowledge of Prometheus, Grafana or similar tools Experience in a platform engineering or SRE environment Certifications in GCP are advantageous What's On Offer Salary up to £80,000 Flexible hybrid Generous holiday allowance Pension scheme Training and certification budget Opportunity to shape a growing cloud platform and influence technical direction Eligo Recruitment is acting as an Employment Business in relation to this vacancy. Eligo is proud to be an equal opportunity employer dedicated to fostering diversity and creating an inclusive and equitable environment for employees and applicants. We actively celebrate and embrace differences, including but not limited to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran status, and disability. We encourage applications from individuals of all backgrounds and experiences and all will be considered for employment without discrimination. At Eligo Recruitment diversity, equity and inclusion is integral to achieving our mission to ensure every workplace reflects the richness of human diversity.
25/07/2026
Full time
GCP Platform Engineer Location: Hybrid / Manchester Salary: Up to £80,000 Benefits Type: Permanent We're partnering with an innovative platform business that is investing heavily in its cloud infrastructure and engineering capabilities. They are looking for an experienced GCP Platform Engineer to help build, automate and scale a modern cloud platform used across multiple products and teams. This is an opportunity to join a collaborative engineering environment where infrastructure is treated as code, automation is a priority, and engineers are empowered to drive technical decisions. The Role As a GCP Platform Engineer, you will be responsible for designing, building and maintaining the company's cloud platform on Google Cloud Platform (GCP). You'll work closely with software engineers and DevOps teams to create scalable, secure and highly available infrastructure. Key responsibilities Designing and implementing cloud infrastructure on GCP Building and maintaining Infrastructure as Code using Terraform Automating infrastructure provisioning and deployment pipelines Managing Kubernetes and containerised workloads Implementing monitoring, logging and observability solutions Driving platform reliability, security and best practices Collaborating with engineering teams to improve developer experience Skills & Experience Essential: Strong commercial experience with Google Cloud Platform (GCP) Extensive experience with Terraform and Infrastructure as Code Experience building CI/CD pipelines Knowledge of Kubernetes / GKE and container technologies Experience with Linux and scripting (Bash, Python or Go) Understanding of networking, IAM and cloud security principles Experience with monitoring and observability tooling Desirable: Experience with GitOps practices Knowledge of Prometheus, Grafana or similar tools Experience in a platform engineering or SRE environment Certifications in GCP are advantageous What's On Offer Salary up to £80,000 Flexible hybrid Generous holiday allowance Pension scheme Training and certification budget Opportunity to shape a growing cloud platform and influence technical direction Eligo Recruitment is acting as an Employment Business in relation to this vacancy. Eligo is proud to be an equal opportunity employer dedicated to fostering diversity and creating an inclusive and equitable environment for employees and applicants. We actively celebrate and embrace differences, including but not limited to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran status, and disability. We encourage applications from individuals of all backgrounds and experiences and all will be considered for employment without discrimination. At Eligo Recruitment diversity, equity and inclusion is integral to achieving our mission to ensure every workplace reflects the richness of human diversity.
The position works closely with IAM architects, business stakeholders and technology partners. In this role, you will: Build and maintain cloud-native platform services supporting the IDAM 2.0 programme. Design, deploy and manage Kubernetes clusters and supporting platform components. Develop Infrastructure as Code (IaC) for repeatable, automated deployments. Implement and maintain CI/CD pipelines and GitOps deployment workflows. Manage cloud networking, connectivity and platform security. Implement platform observability including logging, monitoring, metrics and distributed tracing. Automate platform provisioning, configuration management and operational tasks. Support deployment and operation of identity platform components and supporting services. Implement secure secrets management and certificate lifecycle automation. Configure service mesh technologies and secure service-to-service communication. Manage ingress, API gateways and network routing components. Implement platform resilience, backup, disaster recovery and high availability capabilities. Support workload identity and platform authentication mechanisms. Ensure platform compliance with enterprise security, governance and regulatory requirements. Optimise platform performance, scalability and cost efficiency. Develop operational tooling, scripts and self-service capabilities. Support incident management, troubleshooting and root cause analysis. Collaborate with Security, Architecture and Engineering teams to deliver secure platform services. Contribute to platform standards, operational documentation and runbooks. Support environment management across development, test, pilot and production environments. Drive continuous improvement through automation and platform engineering best practices. To be successful in this role, you should meet the following requirements: Key Skills & Experience Essential Skills Cloud Platform Engineering (GCP) Kubernetes Administration Infrastructure as Code (Terraform or equivalent) CI/CD Pipelines GitOps Linux Administration Networking and Load Balancing Service Mesh Technologies (Istio/Envoy) Container Platforms Secrets Management PKI and Certificate Management Workload Identity Observability (Logging, Monitoring and Tracing) Scripting and Automation DevSecOps Site Reliability Engineering (SRE) Performance and Capacity Management Operational Support Global Deployment Strategies Positive can-do attitude adept at solutionising in large, complex organisations GCS is acting as an Employment Business in relation to this vacancy.
25/07/2026
Contractor
The position works closely with IAM architects, business stakeholders and technology partners. In this role, you will: Build and maintain cloud-native platform services supporting the IDAM 2.0 programme. Design, deploy and manage Kubernetes clusters and supporting platform components. Develop Infrastructure as Code (IaC) for repeatable, automated deployments. Implement and maintain CI/CD pipelines and GitOps deployment workflows. Manage cloud networking, connectivity and platform security. Implement platform observability including logging, monitoring, metrics and distributed tracing. Automate platform provisioning, configuration management and operational tasks. Support deployment and operation of identity platform components and supporting services. Implement secure secrets management and certificate lifecycle automation. Configure service mesh technologies and secure service-to-service communication. Manage ingress, API gateways and network routing components. Implement platform resilience, backup, disaster recovery and high availability capabilities. Support workload identity and platform authentication mechanisms. Ensure platform compliance with enterprise security, governance and regulatory requirements. Optimise platform performance, scalability and cost efficiency. Develop operational tooling, scripts and self-service capabilities. Support incident management, troubleshooting and root cause analysis. Collaborate with Security, Architecture and Engineering teams to deliver secure platform services. Contribute to platform standards, operational documentation and runbooks. Support environment management across development, test, pilot and production environments. Drive continuous improvement through automation and platform engineering best practices. To be successful in this role, you should meet the following requirements: Key Skills & Experience Essential Skills Cloud Platform Engineering (GCP) Kubernetes Administration Infrastructure as Code (Terraform or equivalent) CI/CD Pipelines GitOps Linux Administration Networking and Load Balancing Service Mesh Technologies (Istio/Envoy) Container Platforms Secrets Management PKI and Certificate Management Workload Identity Observability (Logging, Monitoring and Tracing) Scripting and Automation DevSecOps Site Reliability Engineering (SRE) Performance and Capacity Management Operational Support Global Deployment Strategies Positive can-do attitude adept at solutionising in large, complex organisations GCS is acting as an Employment Business in relation to this vacancy.
Cloud & Data Engineering We're Hitachi Digital Services, a global digital solutions and transformation business. Mandate Design, build, and enforce operational excellence of the organisation's cloud platform across private and public cloud environments. Mandatory Skills Platform Engineering for on premises & cloud, IaC, CI/CD pipeline & tools, security, network, landing zones, observability, self servicing, DB automation Key Responsibilities Architect and implement IaC based cloud provisioning using Terraform, CloudFormation, or Pulumi across multi cloud and private cloud environments Design and automate end to end infrastructure lifecycle management: provisioning, patching, updates, scaling, and decommissioning Build and maintain cloud governance guardrails using policy as code (OPA, Sentinel, Azure Policy) to enforce security, cost, and compliance standards Implement FinOps practices including automated rightsizing, reserved instance management, waste detection, and cost allocation tagging Lead CI/CD pipeline engineering for infrastructure deployments, integrating automated testing, security scanning, and approval gates Design self service infrastructure portals enabling development teams to provision compliant environments without manual intervention Mentor and coach 2 Cloud Platform Engineers, establish coding standards, conduct architecture reviews, and drive knowledge sharing Collaborate with SRE and Security teams to embed reliability, observability, and compliance into the platform from the ground up Produce architectural decision records (ADRs), runbooks, and platform documentation for operational excellence Technical Skills & Expertise Expert level IaC proficiency: Terraform (strongly preferred), CloudFormation, Pulumi, or Bicep Deep experience with at least two major cloud platforms: AWS, Azure, GCP, or VMware private cloud Strong experience with Kubernetes (EKS/AKS/GKE), container orchestration, and platform services Configuration management expertise: Ansible, Chef, Puppet, or SaltStack CI/CD tooling: Jenkins, GitLab CI, GitHub Actions, Azure DevOps for infrastructure pipelines Policy as code frameworks: OPA/Rego, HashiCorp Sentinel, Azure Policy, AWS Config Rules Monitoring and observability: Prometheus, Grafana, Datadog, CloudWatch, or Dynatrace Networking fundamentals: VPC/VNet design, load balancers, DNS, CDN, and hybrid connectivity Soft skills & Competencies Strong leadership and mentoring ability - coaches and develops junior engineers Excellent stakeholder management and communication skills across all levels Strategic thinker who balances technical depth with business outcomes Proven ability to drive change, influence without authority, and build consensus Strong analytical and problem solving mindset with attention to detail Ability to manage competing priorities across multiple workstreams simultaneously Qualifications & Experience 7+ years in cloud/infrastructure engineering with 3+ years in a senior or lead capacity Proven track record of delivering large scale IaC transformation and cloud migration programmes Relevant certifications: AWS Solutions Architect Professional, Azure Solutions Architect Expert, CKA/CKAD, or Terraform Associate/Professional Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience) Desirable / Nice to Have Experience with FinOps frameworks and cloud cost optimisation at scale Background in platform engineering product management and internal developer platforms Experience with GitOps workflows (ArgoCD, Flux) and service mesh (Istio, Linkerd) Equal Opportunity Employment We're proud to say we're an equal opportunity employer and welcome all applicants for employment without regard to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruiting process, please let us know so that we can do our best to support your success.
24/07/2026
Full time
Cloud & Data Engineering We're Hitachi Digital Services, a global digital solutions and transformation business. Mandate Design, build, and enforce operational excellence of the organisation's cloud platform across private and public cloud environments. Mandatory Skills Platform Engineering for on premises & cloud, IaC, CI/CD pipeline & tools, security, network, landing zones, observability, self servicing, DB automation Key Responsibilities Architect and implement IaC based cloud provisioning using Terraform, CloudFormation, or Pulumi across multi cloud and private cloud environments Design and automate end to end infrastructure lifecycle management: provisioning, patching, updates, scaling, and decommissioning Build and maintain cloud governance guardrails using policy as code (OPA, Sentinel, Azure Policy) to enforce security, cost, and compliance standards Implement FinOps practices including automated rightsizing, reserved instance management, waste detection, and cost allocation tagging Lead CI/CD pipeline engineering for infrastructure deployments, integrating automated testing, security scanning, and approval gates Design self service infrastructure portals enabling development teams to provision compliant environments without manual intervention Mentor and coach 2 Cloud Platform Engineers, establish coding standards, conduct architecture reviews, and drive knowledge sharing Collaborate with SRE and Security teams to embed reliability, observability, and compliance into the platform from the ground up Produce architectural decision records (ADRs), runbooks, and platform documentation for operational excellence Technical Skills & Expertise Expert level IaC proficiency: Terraform (strongly preferred), CloudFormation, Pulumi, or Bicep Deep experience with at least two major cloud platforms: AWS, Azure, GCP, or VMware private cloud Strong experience with Kubernetes (EKS/AKS/GKE), container orchestration, and platform services Configuration management expertise: Ansible, Chef, Puppet, or SaltStack CI/CD tooling: Jenkins, GitLab CI, GitHub Actions, Azure DevOps for infrastructure pipelines Policy as code frameworks: OPA/Rego, HashiCorp Sentinel, Azure Policy, AWS Config Rules Monitoring and observability: Prometheus, Grafana, Datadog, CloudWatch, or Dynatrace Networking fundamentals: VPC/VNet design, load balancers, DNS, CDN, and hybrid connectivity Soft skills & Competencies Strong leadership and mentoring ability - coaches and develops junior engineers Excellent stakeholder management and communication skills across all levels Strategic thinker who balances technical depth with business outcomes Proven ability to drive change, influence without authority, and build consensus Strong analytical and problem solving mindset with attention to detail Ability to manage competing priorities across multiple workstreams simultaneously Qualifications & Experience 7+ years in cloud/infrastructure engineering with 3+ years in a senior or lead capacity Proven track record of delivering large scale IaC transformation and cloud migration programmes Relevant certifications: AWS Solutions Architect Professional, Azure Solutions Architect Expert, CKA/CKAD, or Terraform Associate/Professional Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience) Desirable / Nice to Have Experience with FinOps frameworks and cloud cost optimisation at scale Background in platform engineering product management and internal developer platforms Experience with GitOps workflows (ArgoCD, Flux) and service mesh (Istio, Linkerd) Equal Opportunity Employment We're proud to say we're an equal opportunity employer and welcome all applicants for employment without regard to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruiting process, please let us know so that we can do our best to support your success.
We're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people centric and here to power good. Every day, we future proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our company and customers from what's now to what's next. We make it happen through the power of acceleration. Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us. Job Description Mandatory Skills: Plat Engineering for on premises & cloud, IaC, CI/CD pipeline & tools, security, network, landing zones, observability, self servicing, DB automation ROLE PURPOSE Lead the design, build, and operational excellence of the organisation's cloud platform across private and public cloud environments. Own the end to end IaC based provisioning, configuration management, patch orchestration, and cloud governance framework. Drive the transformation from manual infrastructure operations to a fully automated, self service, policy governed cloud platform that delivers speed, security, cost efficiency, and reliability. KEY RESPONSIBILITIES Architect and implement IaC based cloud provisioning using Terraform, CloudFormation, or Pulumi across multi cloud and private cloud environments Design and automate end to end infrastructure lifecycle management: provisioning, patching, updates, scaling, and decommissioning Build and maintain cloud governance guardrails using policy as code (OPA, Sentinel, Azure Policy) to enforce security, cost, and compliance standards Implement FinOps practices including automated rightsizing, reserved instance management, waste detection, and cost allocation tagging Lead CI/CD pipeline engineering for infrastructure deployments, integrating automated testing, security scanning, and approval gates Design self service infrastructure portals enabling development teams to provision compliant environments without manual intervention Mentor and coach 2 Cloud Platform Engineers, establish coding standards, conduct architecture reviews, and drive knowledge sharing Collaborate with SRE and Security teams to embed reliability, observability, and compliance into the platform from the ground up Produce architectural decision records (ADRs), runbooks, and platform documentation for operational excellence TECHNICAL SKILLS & EXPERTISE Expert level IaC proficiency: Terraform (strongly preferred), CloudFormation, Pulumi, or Bicep Deep experience with at least two major cloud platforms: AWS, Azure, GCP, or VMware private cloud Strong experience with Kubernetes (EKS/AKS/GKE), container orchestration, and platform services Configuration management expertise: Ansible, Chef, Puppet, or SaltStack CI/CD tooling: Jenkins, GitLab CI, GitHub Actions, Azure DevOps for infrastructure pipelines Policy as code frameworks: OPA/Rego, HashiCorp Sentinel, Azure Policy, AWS Config Rules Monitoring and observability: Prometheus, Grafana, Datadog, CloudWatch, or Dynatrace Networking fundamentals: VPC/VNet design, load balancers, DNS, CDN, and hybrid connectivity SOFT SKILLS & COMPETENCIES Strong leadership and mentoring ability - coaches and develops junior engineers Excellent stakeholder management and communication skills across all levels Strategic thinker who balances technical depth with business outcomes Proven ability to drive change, influence without authority, and build consensus Strong analytical and problem solving mindset with attention to detail Ability to manage competing priorities across multiple workstreams simultaneously QUALIFICATIONS & EXPERIENCE 7+ years in cloud/infrastructure engineering with 3+ years in a senior or lead capacity Proven track record of delivering large scale IaC transformation and cloud migration programmes Relevant certifications: AWS Solutions Architect Professional, Azure Solutions Architect Expert, CKA/CKAD, or Terraform Associate/Professional Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience) DESIRABLE / NICE TO HAVE Experience with FinOps frameworks and cloud cost optimisation at scale Background in platform engineering product management and internal developer platforms Experience with GitOps workflows (ArgoCD, Flux) and service mesh (Istio, Linkerd) About Us We're a global, team of innovators. Together, we harness engineering excellence and passion to co create meaningful solutions to complex challenges. We turn organisations into data driven leaders that can make a positive impact on their industries and society. If you believe that innovation can bring a better tomorrow closer to today, this is the place for you. Fostering innovation through diverse perspectives Hitachi is a global company operating across a wide range of industries and regions. One of the things that sets Hitachi apart is the diversity of our business and people, which drives our innovation and growth. We are committed to building an inclusive culture based on mutual respect and merit based systems. We believe that when people feel valued, heard, and safe to express themselves, they do their best work. How we look after you We help take care of your today and tomorrow with industry leading benefits, support, and services that look after your holistic health and wellbeing. We're also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We're always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you'll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with. We're proud to say we're an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
24/07/2026
Full time
We're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people centric and here to power good. Every day, we future proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our company and customers from what's now to what's next. We make it happen through the power of acceleration. Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us. Job Description Mandatory Skills: Plat Engineering for on premises & cloud, IaC, CI/CD pipeline & tools, security, network, landing zones, observability, self servicing, DB automation ROLE PURPOSE Lead the design, build, and operational excellence of the organisation's cloud platform across private and public cloud environments. Own the end to end IaC based provisioning, configuration management, patch orchestration, and cloud governance framework. Drive the transformation from manual infrastructure operations to a fully automated, self service, policy governed cloud platform that delivers speed, security, cost efficiency, and reliability. KEY RESPONSIBILITIES Architect and implement IaC based cloud provisioning using Terraform, CloudFormation, or Pulumi across multi cloud and private cloud environments Design and automate end to end infrastructure lifecycle management: provisioning, patching, updates, scaling, and decommissioning Build and maintain cloud governance guardrails using policy as code (OPA, Sentinel, Azure Policy) to enforce security, cost, and compliance standards Implement FinOps practices including automated rightsizing, reserved instance management, waste detection, and cost allocation tagging Lead CI/CD pipeline engineering for infrastructure deployments, integrating automated testing, security scanning, and approval gates Design self service infrastructure portals enabling development teams to provision compliant environments without manual intervention Mentor and coach 2 Cloud Platform Engineers, establish coding standards, conduct architecture reviews, and drive knowledge sharing Collaborate with SRE and Security teams to embed reliability, observability, and compliance into the platform from the ground up Produce architectural decision records (ADRs), runbooks, and platform documentation for operational excellence TECHNICAL SKILLS & EXPERTISE Expert level IaC proficiency: Terraform (strongly preferred), CloudFormation, Pulumi, or Bicep Deep experience with at least two major cloud platforms: AWS, Azure, GCP, or VMware private cloud Strong experience with Kubernetes (EKS/AKS/GKE), container orchestration, and platform services Configuration management expertise: Ansible, Chef, Puppet, or SaltStack CI/CD tooling: Jenkins, GitLab CI, GitHub Actions, Azure DevOps for infrastructure pipelines Policy as code frameworks: OPA/Rego, HashiCorp Sentinel, Azure Policy, AWS Config Rules Monitoring and observability: Prometheus, Grafana, Datadog, CloudWatch, or Dynatrace Networking fundamentals: VPC/VNet design, load balancers, DNS, CDN, and hybrid connectivity SOFT SKILLS & COMPETENCIES Strong leadership and mentoring ability - coaches and develops junior engineers Excellent stakeholder management and communication skills across all levels Strategic thinker who balances technical depth with business outcomes Proven ability to drive change, influence without authority, and build consensus Strong analytical and problem solving mindset with attention to detail Ability to manage competing priorities across multiple workstreams simultaneously QUALIFICATIONS & EXPERIENCE 7+ years in cloud/infrastructure engineering with 3+ years in a senior or lead capacity Proven track record of delivering large scale IaC transformation and cloud migration programmes Relevant certifications: AWS Solutions Architect Professional, Azure Solutions Architect Expert, CKA/CKAD, or Terraform Associate/Professional Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience) DESIRABLE / NICE TO HAVE Experience with FinOps frameworks and cloud cost optimisation at scale Background in platform engineering product management and internal developer platforms Experience with GitOps workflows (ArgoCD, Flux) and service mesh (Istio, Linkerd) About Us We're a global, team of innovators. Together, we harness engineering excellence and passion to co create meaningful solutions to complex challenges. We turn organisations into data driven leaders that can make a positive impact on their industries and society. If you believe that innovation can bring a better tomorrow closer to today, this is the place for you. Fostering innovation through diverse perspectives Hitachi is a global company operating across a wide range of industries and regions. One of the things that sets Hitachi apart is the diversity of our business and people, which drives our innovation and growth. We are committed to building an inclusive culture based on mutual respect and merit based systems. We believe that when people feel valued, heard, and safe to express themselves, they do their best work. How we look after you We help take care of your today and tomorrow with industry leading benefits, support, and services that look after your holistic health and wellbeing. We're also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We're always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you'll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with. We're proud to say we're an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
Function Cloud & Data Engineering About the Company We're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people centric and here to power good. Every day, we future prove urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our company and customers from what's now to what's next. We make it happen through the power of acceleration. Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us. Job Description Mandatory Skills Plat Engineering for on premises & cloud, IaC, CI/CD pipeline & tools, security, network, landing zones, observability, self servicing, DB automation Role Purpose Lead the design, build, and operational excellence of the organisation's cloud platform across private and public cloud environments. Own the end to end IaC based provisioning, configuration management, patch orchestration, and cloud governance framework. Drive the transformation from manual infrastructure operations to a fully automated, self service, policy governed cloud platform that delivers speed, security, cost efficiency, and reliability. Key Responsibilities Architect and implement IaC based cloud provisioning using Terraform, CloudFormation, or Pulumi across multi cloud and private cloud environments Design and automate end to end infrastructure lifecycle management: provisioning, patching, updates, scaling, and decommissioning Build and maintain cloud governance guardrails using policy as code (OPA, Sentinel, Azure Policy) to enforce security, cost, and compliance standards Implement FinOps practices including automated rightsizing, reserved instance management, waste detection, and cost allocation tagging Lead CI/CD pipeline engineering for infrastructure deployments, integrating automated testing, security scanning, and approval gates Design self service infrastructure portals enabling development teams to provision compliant environments without manual intervention Mentor and coach 2 Cloud Platform Engineers, establish coding standards, conduct architecture reviews, and drive knowledge sharing Collaborate with SRE and Security teams to embed reliability, observability, and compliance into the platform from the ground up Produce architectural decision records (ADRs), runbooks, and platform documentation for operational excellence Technical Skills & Expertise Expert level IaC proficiency: Terraform (strongly preferred), CloudFormation, Pulumi, or Bicep Deep experience with at least two major cloud platforms: AWS, Azure, GCP, or VMware private cloud Strong experience with Kubernetes (EKS/AKS/GKE), container orchestration, and platform services Configuration management expertise: Ansible, Chef, Puppet, or SaltStack CI/CD tooling: Jenkins, GitLab CI, GitHub Actions, Azure DevOps for infrastructure pipelines Policy as code frameworks: OPA/Rego, HashiCorp Sentinel, Azure Policy, AWS Config Rules Monitoring and observability: Prometheus, Grafana, Datadog, CloudWatch, or Dynatrace Networking fundamentals: VPC/VNet design, load balancers, DNS, CDN, and hybrid connectivity Soft Skills & Competencies Strong leadership and mentoring ability - coaches and develops junior engineers Excellent stakeholder management and communication skills across all levels Strategic thinker who balances technical depth with business outcomes Proven ability to drive change, influence without authority, and build consensus Strong analytical and problem solving mindset with attention to detail Ability to manage competing priorities across multiple workstreams simultaneously Qualifications & Experience 7+ years in cloud/infrastructure engineering with 3+ years in a senior or lead capacity Proven track record of delivering large scale IaC transformation and cloud migration programmes Relevant certifications: AWS Solutions Architect Professional, Azure Solutions Architect Expert, CKA/CKAD, or Terraform Associate/Professional Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience) Desirable / Nice to Have Experience with FinOps frameworks and cloud cost optimisation at scale Background in platform engineering product management and internal developer platforms Experience with GitOps workflows (ArgoCD, Flux) and service mesh (Istio, Linkerd) Benefits & Well being We help take care of your today and tomorrow with industry leading benefits, support, and services that look after your holistic health and wellbeing. We're also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We're always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you'll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with. Equal Opportunity Statement We're proud to say we're an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
24/07/2026
Full time
Function Cloud & Data Engineering About the Company We're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people centric and here to power good. Every day, we future prove urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our company and customers from what's now to what's next. We make it happen through the power of acceleration. Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us. Job Description Mandatory Skills Plat Engineering for on premises & cloud, IaC, CI/CD pipeline & tools, security, network, landing zones, observability, self servicing, DB automation Role Purpose Lead the design, build, and operational excellence of the organisation's cloud platform across private and public cloud environments. Own the end to end IaC based provisioning, configuration management, patch orchestration, and cloud governance framework. Drive the transformation from manual infrastructure operations to a fully automated, self service, policy governed cloud platform that delivers speed, security, cost efficiency, and reliability. Key Responsibilities Architect and implement IaC based cloud provisioning using Terraform, CloudFormation, or Pulumi across multi cloud and private cloud environments Design and automate end to end infrastructure lifecycle management: provisioning, patching, updates, scaling, and decommissioning Build and maintain cloud governance guardrails using policy as code (OPA, Sentinel, Azure Policy) to enforce security, cost, and compliance standards Implement FinOps practices including automated rightsizing, reserved instance management, waste detection, and cost allocation tagging Lead CI/CD pipeline engineering for infrastructure deployments, integrating automated testing, security scanning, and approval gates Design self service infrastructure portals enabling development teams to provision compliant environments without manual intervention Mentor and coach 2 Cloud Platform Engineers, establish coding standards, conduct architecture reviews, and drive knowledge sharing Collaborate with SRE and Security teams to embed reliability, observability, and compliance into the platform from the ground up Produce architectural decision records (ADRs), runbooks, and platform documentation for operational excellence Technical Skills & Expertise Expert level IaC proficiency: Terraform (strongly preferred), CloudFormation, Pulumi, or Bicep Deep experience with at least two major cloud platforms: AWS, Azure, GCP, or VMware private cloud Strong experience with Kubernetes (EKS/AKS/GKE), container orchestration, and platform services Configuration management expertise: Ansible, Chef, Puppet, or SaltStack CI/CD tooling: Jenkins, GitLab CI, GitHub Actions, Azure DevOps for infrastructure pipelines Policy as code frameworks: OPA/Rego, HashiCorp Sentinel, Azure Policy, AWS Config Rules Monitoring and observability: Prometheus, Grafana, Datadog, CloudWatch, or Dynatrace Networking fundamentals: VPC/VNet design, load balancers, DNS, CDN, and hybrid connectivity Soft Skills & Competencies Strong leadership and mentoring ability - coaches and develops junior engineers Excellent stakeholder management and communication skills across all levels Strategic thinker who balances technical depth with business outcomes Proven ability to drive change, influence without authority, and build consensus Strong analytical and problem solving mindset with attention to detail Ability to manage competing priorities across multiple workstreams simultaneously Qualifications & Experience 7+ years in cloud/infrastructure engineering with 3+ years in a senior or lead capacity Proven track record of delivering large scale IaC transformation and cloud migration programmes Relevant certifications: AWS Solutions Architect Professional, Azure Solutions Architect Expert, CKA/CKAD, or Terraform Associate/Professional Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience) Desirable / Nice to Have Experience with FinOps frameworks and cloud cost optimisation at scale Background in platform engineering product management and internal developer platforms Experience with GitOps workflows (ArgoCD, Flux) and service mesh (Istio, Linkerd) Benefits & Well being We help take care of your today and tomorrow with industry leading benefits, support, and services that look after your holistic health and wellbeing. We're also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We're always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you'll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with. Equal Opportunity Statement We're proud to say we're an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
We Are Skyral: We believe every decision maker can be empowered by technology. Skyral combines AI, leading edge simulation technology and world class expertise to transform the decision making experience. Our products and services enable faster and more confident decisions in a complex, unforgiving world. We deploy practical, intuitive and efficient solutions to governments and enterprises, delivering outstanding outcomes at the speed of relevance. At Omnia Training, we've brought together some of the UK's most innovative defence training organisations under one powerful mission: to transform the British Army's training system and create the best-trained Army in the world. Omnia Training is redefining the British Army's collective training. To do that, we are looking for the best and brightest minds from across the UK. Omnia Training is at the heart of the UK's bold Land Industrial Strategy. This is more than a job - it's a mission. You will be part of a high-impact, collaborative environment, where every person in our team plays a critical role in delivering Omnia Training's vision; designing, delivering, and transforming collective training. Please note that this role will require onsite working in Warminster. Due to the nature of this role, Skyral can only consider applications from candidates who live in the UK and are eligible for SC clearance. What You'll Be Responsible For: Deploy and configure systems on MODCloud / D2S / OpenShift Manage Kubernetes environments, networking and connectivity, as well as security-aligned configurations. Build and maintain CI/CD pipelines, GitOps workflows, and environment configurations to ensure repeatable and reliable deployments. Support running systems by monitoring health, diagnosing platform and infrastructure issues and resolving deployment failures. Work closely with integration engineers, customers and modelling engineers. Support and maintain existing simulation and training systems, as well as existing deployment and virtualisation tools. Apply SRE practices to improve system reliability, including observability (metrics, logs, tracing), incident response, and root cause analysis. What We Are Looking For: This is not a pure cloud or greenfield platform role. You will be working across cloud-native services, legacy systems, and integrated simulation environments, ensuring they operate reliably as a single platform. Strong systems and infrastructure mindset Comfortable working across modern and legacy environments Calm under pressure during outages or failures Pragmatic and delivery-focused, with a bias toward keeping systems running. Strong collaborator across engineering disciplines Adopts an SRE mindset, focusing on reliability, observability, and continuous improvement of running systems. Key Technical Proficiencies: Expert working knowledge of Kubernetes, Helm, Teraform, Ansible, and Docker. Understanding of Distributed Systems in production. Experience working within constrained or regulated environments (e.g. MODCloud, D2S, OpenShift) and adapting to their tooling and limitations. Experience building and operating CI/CD pipelines with automated deployment workflows. Familiarity with GitOps approaches and tools such as ArgoCD. Strong understanding of Network fundamentals, Zero Trust solutions, service to service communications and distributed system connectivity. Ability to diagnose issues across infrastructure, networking, and application layers. Experience supporting or integrating legacy and non-cloud-native systems alongside modern infrastructure. Experience in developing with Go or Python as well as shell scripting. Experience applying Site Reliability Engineering (SRE) practices such as monitoring, alerting, incident response, and service reliability improvement. Note: Please feel empowered to apply for this position, even if you think you may only align with some of the qualities listed above. Your unique skills and perspectives could be just what we're looking for. What We Can Offer You: Unlimited Paid Holiday - we value and support the need to maintain a strong work-life balance. Hybrid Working - we understand that a one-size-fits all approach doesn't suit everyone. Flexible Working Hours - We're not bound by the 9-to-5 model. Collaborate with your manager on determining a work schedule that suits you. Enhanced Parental Leave - we're proud to offer 26 weeks maternity leave and 4 weeks paternity leave at full pay. Private Medical & Dental Insurance - offered through Bupa. Honest about Compensation - We maintain a well defined salary range which a member of the Talent Team will discuss with you during the first call. Healthy Snacks & Drinks Provided - If you decide to come into the office, we have a range of snacks and drinks for you to enjoy. At Skyral, we are committed to fostering a culture of diversity, equality and inclusion. We also ensure that individuals with disabilities have access to reasonable adjustments. If you require such accommodations during the job application process we ask that you inform a member of our Talent Team.
24/07/2026
Full time
We Are Skyral: We believe every decision maker can be empowered by technology. Skyral combines AI, leading edge simulation technology and world class expertise to transform the decision making experience. Our products and services enable faster and more confident decisions in a complex, unforgiving world. We deploy practical, intuitive and efficient solutions to governments and enterprises, delivering outstanding outcomes at the speed of relevance. At Omnia Training, we've brought together some of the UK's most innovative defence training organisations under one powerful mission: to transform the British Army's training system and create the best-trained Army in the world. Omnia Training is redefining the British Army's collective training. To do that, we are looking for the best and brightest minds from across the UK. Omnia Training is at the heart of the UK's bold Land Industrial Strategy. This is more than a job - it's a mission. You will be part of a high-impact, collaborative environment, where every person in our team plays a critical role in delivering Omnia Training's vision; designing, delivering, and transforming collective training. Please note that this role will require onsite working in Warminster. Due to the nature of this role, Skyral can only consider applications from candidates who live in the UK and are eligible for SC clearance. What You'll Be Responsible For: Deploy and configure systems on MODCloud / D2S / OpenShift Manage Kubernetes environments, networking and connectivity, as well as security-aligned configurations. Build and maintain CI/CD pipelines, GitOps workflows, and environment configurations to ensure repeatable and reliable deployments. Support running systems by monitoring health, diagnosing platform and infrastructure issues and resolving deployment failures. Work closely with integration engineers, customers and modelling engineers. Support and maintain existing simulation and training systems, as well as existing deployment and virtualisation tools. Apply SRE practices to improve system reliability, including observability (metrics, logs, tracing), incident response, and root cause analysis. What We Are Looking For: This is not a pure cloud or greenfield platform role. You will be working across cloud-native services, legacy systems, and integrated simulation environments, ensuring they operate reliably as a single platform. Strong systems and infrastructure mindset Comfortable working across modern and legacy environments Calm under pressure during outages or failures Pragmatic and delivery-focused, with a bias toward keeping systems running. Strong collaborator across engineering disciplines Adopts an SRE mindset, focusing on reliability, observability, and continuous improvement of running systems. Key Technical Proficiencies: Expert working knowledge of Kubernetes, Helm, Teraform, Ansible, and Docker. Understanding of Distributed Systems in production. Experience working within constrained or regulated environments (e.g. MODCloud, D2S, OpenShift) and adapting to their tooling and limitations. Experience building and operating CI/CD pipelines with automated deployment workflows. Familiarity with GitOps approaches and tools such as ArgoCD. Strong understanding of Network fundamentals, Zero Trust solutions, service to service communications and distributed system connectivity. Ability to diagnose issues across infrastructure, networking, and application layers. Experience supporting or integrating legacy and non-cloud-native systems alongside modern infrastructure. Experience in developing with Go or Python as well as shell scripting. Experience applying Site Reliability Engineering (SRE) practices such as monitoring, alerting, incident response, and service reliability improvement. Note: Please feel empowered to apply for this position, even if you think you may only align with some of the qualities listed above. Your unique skills and perspectives could be just what we're looking for. What We Can Offer You: Unlimited Paid Holiday - we value and support the need to maintain a strong work-life balance. Hybrid Working - we understand that a one-size-fits all approach doesn't suit everyone. Flexible Working Hours - We're not bound by the 9-to-5 model. Collaborate with your manager on determining a work schedule that suits you. Enhanced Parental Leave - we're proud to offer 26 weeks maternity leave and 4 weeks paternity leave at full pay. Private Medical & Dental Insurance - offered through Bupa. Honest about Compensation - We maintain a well defined salary range which a member of the Talent Team will discuss with you during the first call. Healthy Snacks & Drinks Provided - If you decide to come into the office, we have a range of snacks and drinks for you to enjoy. At Skyral, we are committed to fostering a culture of diversity, equality and inclusion. We also ensure that individuals with disabilities have access to reasonable adjustments. If you require such accommodations during the job application process we ask that you inform a member of our Talent Team.
Position Overview At CreateFuture, the Senior Cloud Engineer is a vital individual contributor and a technical driving force within our delivery teams. You will work closely with clients and delivery teams to bridge the implementation gap, assisting with architectural designs and transforming them into production systems using autonomous and AI-native platforms. This role requires a modern, hybrid engineering approach, blending infrastructure mastery with software development and data engineering principles. You will act as a technical expert on projects, building trusted relationships with client stakeholders, establishing engineering best practices, and mentoring engineers within the Cloud capability. Furthermore, you will actively contribute to setting engineering standards that influence the whole CreateFuture organisation. Key Responsibilities High-Quality Execution: Proactively contribute to project teams by providing technical guidance, resolving complex technical blockages, supporting migration efforts and writing high-quality, spec-driven code. Platform Engineering: Build and maintain independent, self-service developer platforms that allow engineering teams to spin up compliant, ephemer environments automatically using GitOps workflows. AI & Agentic Infrastructure: Implement the data pipelines, workflow orchestrations, and specialised compute footprints needed to support enterprise AI applications, using technologies like AWS Bedrock and AWS AgentCore. Observability & Reliability: Build robust monitoring and observability pipelines to ensure the health, performance, and security of distributed cloud applications and AI models. FinOps Standards: Embed automated cost-estimation tools into the standard CI/CD deployment pipelines, ensuring auto-shutdown policies and resource optimisation are configured by default. Consultancy & Client Alignment Client Advisory: Collaborate with diverse stakeholders, including tech leads, project managers, and enterprise architects, to ensure successful project delivery according to timelines and budgets. Commercial Growth: Have a growth mindset on projects, recognising potential operational bottlenecks or new requirements, and own communicating these throughout the team for new opportunities. Best Practice Champion: Establish best practice tools and processes within client environments, confidently challenging legacy ways of working where appropriate. Capability & Leadership Support Mentoring: Support and mentor engineers within CreateFuture and our client teams, helping them navigate their learning and technical skill gaps. Community Contribution: Actively engage in the internal engineering communities, helping to build out an internal repository of reference architectures, Infrastructure as Code templates, and technical blogs. Culture and Feedback: Help foster an inclusive, collaborative, and psychologically safe team environment by providing timely, open, and honest technical feedback. Skills & Experience Core Technical Capabilities Candidates must demonstrate a hybrid balance across the three strategic pillars: Software Engineering (30%), Infrastructure (40%), and Data Engineering (30%). Infrastructure & Automation (40%): Lots of experience in writing declarative Infrastructure as Code using Terraform, container orchestration with Kubernetes (managing clusters, nodes, and pods), and building CI/CD pipelines via GitHub Actions. Software & Agentic Engineering (30%): Proficiency with modern scripting languages for building custom AI agents, configuring workflow orchestration, connecting enterprise API integrations, and implementing agentic guardrails. Data Engineering (30%): Practical experience implementing components of data pipelines, including real-time streaming tools (AWS Kinesis, Kafka), data orchestration (dbt, Airflow), and managing vector databases for RAG architectures. Observability & Cost Management: SRE/Platform experience with the practical application of real-time monitoring and cloud cost optimisation using native CSP tools or utilities like Infracost and Karpenter. Domain & Sector Experience Regulated Industries: Experience delivering secure platforms within highly regulated environments, such as iGaming, financial services, or banking, where zero trust security and strict compliance controls are mandatory is highly advantageous. Consulting: Proven experience operating in a client-facing or consulting engineering role, demonstrating strong communication skills, relationship-building capabilities, and adaptability across changing technical contexts. Preferred Knowledge-level & Certifications Associate-level cloud certifications (e.g., AWS Solutions Architect Associate or Azure AZ-104). At least one professional-level certification. Progressing toward or holding multiple professional certifications such as CKA (Certified Kubernetes Administrator), AWS ML Engineer Associate, or FinOps Certified Practitioner. What we'll offer you: We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You'll be part of a supportive and friendly culture, surrounded by smart, curious people who care deeply about what they do. We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed. We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.
21/07/2026
Full time
Position Overview At CreateFuture, the Senior Cloud Engineer is a vital individual contributor and a technical driving force within our delivery teams. You will work closely with clients and delivery teams to bridge the implementation gap, assisting with architectural designs and transforming them into production systems using autonomous and AI-native platforms. This role requires a modern, hybrid engineering approach, blending infrastructure mastery with software development and data engineering principles. You will act as a technical expert on projects, building trusted relationships with client stakeholders, establishing engineering best practices, and mentoring engineers within the Cloud capability. Furthermore, you will actively contribute to setting engineering standards that influence the whole CreateFuture organisation. Key Responsibilities High-Quality Execution: Proactively contribute to project teams by providing technical guidance, resolving complex technical blockages, supporting migration efforts and writing high-quality, spec-driven code. Platform Engineering: Build and maintain independent, self-service developer platforms that allow engineering teams to spin up compliant, ephemer environments automatically using GitOps workflows. AI & Agentic Infrastructure: Implement the data pipelines, workflow orchestrations, and specialised compute footprints needed to support enterprise AI applications, using technologies like AWS Bedrock and AWS AgentCore. Observability & Reliability: Build robust monitoring and observability pipelines to ensure the health, performance, and security of distributed cloud applications and AI models. FinOps Standards: Embed automated cost-estimation tools into the standard CI/CD deployment pipelines, ensuring auto-shutdown policies and resource optimisation are configured by default. Consultancy & Client Alignment Client Advisory: Collaborate with diverse stakeholders, including tech leads, project managers, and enterprise architects, to ensure successful project delivery according to timelines and budgets. Commercial Growth: Have a growth mindset on projects, recognising potential operational bottlenecks or new requirements, and own communicating these throughout the team for new opportunities. Best Practice Champion: Establish best practice tools and processes within client environments, confidently challenging legacy ways of working where appropriate. Capability & Leadership Support Mentoring: Support and mentor engineers within CreateFuture and our client teams, helping them navigate their learning and technical skill gaps. Community Contribution: Actively engage in the internal engineering communities, helping to build out an internal repository of reference architectures, Infrastructure as Code templates, and technical blogs. Culture and Feedback: Help foster an inclusive, collaborative, and psychologically safe team environment by providing timely, open, and honest technical feedback. Skills & Experience Core Technical Capabilities Candidates must demonstrate a hybrid balance across the three strategic pillars: Software Engineering (30%), Infrastructure (40%), and Data Engineering (30%). Infrastructure & Automation (40%): Lots of experience in writing declarative Infrastructure as Code using Terraform, container orchestration with Kubernetes (managing clusters, nodes, and pods), and building CI/CD pipelines via GitHub Actions. Software & Agentic Engineering (30%): Proficiency with modern scripting languages for building custom AI agents, configuring workflow orchestration, connecting enterprise API integrations, and implementing agentic guardrails. Data Engineering (30%): Practical experience implementing components of data pipelines, including real-time streaming tools (AWS Kinesis, Kafka), data orchestration (dbt, Airflow), and managing vector databases for RAG architectures. Observability & Cost Management: SRE/Platform experience with the practical application of real-time monitoring and cloud cost optimisation using native CSP tools or utilities like Infracost and Karpenter. Domain & Sector Experience Regulated Industries: Experience delivering secure platforms within highly regulated environments, such as iGaming, financial services, or banking, where zero trust security and strict compliance controls are mandatory is highly advantageous. Consulting: Proven experience operating in a client-facing or consulting engineering role, demonstrating strong communication skills, relationship-building capabilities, and adaptability across changing technical contexts. Preferred Knowledge-level & Certifications Associate-level cloud certifications (e.g., AWS Solutions Architect Associate or Azure AZ-104). At least one professional-level certification. Progressing toward or holding multiple professional certifications such as CKA (Certified Kubernetes Administrator), AWS ML Engineer Associate, or FinOps Certified Practitioner. What we'll offer you: We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You'll be part of a supportive and friendly culture, surrounded by smart, curious people who care deeply about what they do. We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed. We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.
Position Overview At CreateFuture, the Senior Cloud Engineer is a vital individual contributor and a technical driving force within our delivery teams. You will work closely with clients and delivery teams to bridge the implementation gap, assisting with architectural designs and transforming them into production systems using autonomous and AI-native platforms. This role requires a modern, hybrid engineering approach, blending infrastructure mastery with software development and data engineering principles. You will act as a technical expert on projects, building trusted relationships with client stakeholders, establishing engineering best practices, and mentoring engineers within the Cloud capability. Furthermore, you will actively contribute to setting engineering standards that influence the whole CreateFuture organisation. Key Responsibilities High-Quality Execution: Proactively contribute to project teams by providing technical guidance, resolving complex technical blockages, supporting migration efforts and writing high-quality, spec-driven code. Platform Engineering: Build and maintain independent, self-service developer platforms that allow engineering teams to spin up compliant, ephemer environments automatically using GitOps workflows. AI & Agentic Infrastructure: Implement the data pipelines, workflow orchestrations, and specialised compute footprints needed to support enterprise AI applications, using technologies like AWS Bedrock and AWS AgentCore. Observability & Reliability: Build robust monitoring and observability pipelines to ensure the health, performance, and security of distributed cloud applications and AI models. FinOps Standards: Embed automated cost-estimation tools into the standard CI/CD deployment pipelines, ensuring auto-shutdown policies and resource optimisation are configured by default. Consultancy & Client Alignment Client Advisory: Collaborate with diverse stakeholders, including tech leads, project managers, and enterprise architects, to ensure successful project delivery according to timelines and budgets. Commercial Growth: Have a growth mindset on projects, recognising potential operational bottlenecks or new requirements, and own communicating these throughout the team for new opportunities. Best Practice Champion: Establish best practice tools and processes within client environments, confidently challenging legacy ways of working where appropriate. Capability & Leadership Support Mentoring: Support and mentor engineers within CreateFuture and our client teams, helping them navigate their learning and technical skill gaps. Community Contribution: Actively engage in the internal engineering communities, helping to build out an internal repository of reference architectures, Infrastructure as Code templates, and technical blogs. Culture and Feedback: Help foster an inclusive, collaborative, and psychologically safe team environment by providing timely, open, and honest technical feedback. Skills & Experience Core Technical Capabilities Candidates must demonstrate a hybrid balance across the three strategic pillars: Software Engineering (30%), Infrastructure (40%), and Data Engineering (30%). Infrastructure & Automation (40%): Lots of experience in writing declarative Infrastructure as Code using Terraform, container orchestration with Kubernetes (managing clusters, nodes, and pods), and building CI/CD pipelines via GitHub Actions. Software & Agentic Engineering (30%): Proficiency with modern scripting languages for building custom AI agents, configuring workflow orchestration, connecting enterprise API integrations, and implementing agentic guardrails. Data Engineering (30%): Practical experience implementing components of data pipelines, including real-time streaming tools (AWS Kinesis, Kafka), data orchestration (dbt, Airflow), and managing vector databases for RAG architectures. Observability & Cost Management: SRE/Platform experience with the practical application of real-time monitoring and cloud cost optimisation using native CSP tools or utilities like Infracost and Karpenter. Domain & Sector Experience Regulated Industries: Experience delivering secure platforms within highly regulated environments, such as iGaming, financial services, or banking, where zero trust security and strict compliance controls are mandatory is highly advantageous. Consulting: Proven experience operating in a client-facing or consulting engineering role, demonstrating strong communication skills, relationship-building capabilities, and adaptability across changing technical contexts. Preferred Knowledge-level & Certifications Associate-level cloud certifications (e.g., AWS Solutions Architect Associate or Azure AZ-104). At least one professional-level certification. Progressing toward or holding multiple professional certifications such as CKA (Certified Kubernetes Administrator), AWS ML Engineer Associate, or FinOps Certified Practitioner. What we'll offer you: We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You'll be part of a supportive and friendly culture, surrounded by smart, curious people who care deeply about what they do. We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed. We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.
21/07/2026
Full time
Position Overview At CreateFuture, the Senior Cloud Engineer is a vital individual contributor and a technical driving force within our delivery teams. You will work closely with clients and delivery teams to bridge the implementation gap, assisting with architectural designs and transforming them into production systems using autonomous and AI-native platforms. This role requires a modern, hybrid engineering approach, blending infrastructure mastery with software development and data engineering principles. You will act as a technical expert on projects, building trusted relationships with client stakeholders, establishing engineering best practices, and mentoring engineers within the Cloud capability. Furthermore, you will actively contribute to setting engineering standards that influence the whole CreateFuture organisation. Key Responsibilities High-Quality Execution: Proactively contribute to project teams by providing technical guidance, resolving complex technical blockages, supporting migration efforts and writing high-quality, spec-driven code. Platform Engineering: Build and maintain independent, self-service developer platforms that allow engineering teams to spin up compliant, ephemer environments automatically using GitOps workflows. AI & Agentic Infrastructure: Implement the data pipelines, workflow orchestrations, and specialised compute footprints needed to support enterprise AI applications, using technologies like AWS Bedrock and AWS AgentCore. Observability & Reliability: Build robust monitoring and observability pipelines to ensure the health, performance, and security of distributed cloud applications and AI models. FinOps Standards: Embed automated cost-estimation tools into the standard CI/CD deployment pipelines, ensuring auto-shutdown policies and resource optimisation are configured by default. Consultancy & Client Alignment Client Advisory: Collaborate with diverse stakeholders, including tech leads, project managers, and enterprise architects, to ensure successful project delivery according to timelines and budgets. Commercial Growth: Have a growth mindset on projects, recognising potential operational bottlenecks or new requirements, and own communicating these throughout the team for new opportunities. Best Practice Champion: Establish best practice tools and processes within client environments, confidently challenging legacy ways of working where appropriate. Capability & Leadership Support Mentoring: Support and mentor engineers within CreateFuture and our client teams, helping them navigate their learning and technical skill gaps. Community Contribution: Actively engage in the internal engineering communities, helping to build out an internal repository of reference architectures, Infrastructure as Code templates, and technical blogs. Culture and Feedback: Help foster an inclusive, collaborative, and psychologically safe team environment by providing timely, open, and honest technical feedback. Skills & Experience Core Technical Capabilities Candidates must demonstrate a hybrid balance across the three strategic pillars: Software Engineering (30%), Infrastructure (40%), and Data Engineering (30%). Infrastructure & Automation (40%): Lots of experience in writing declarative Infrastructure as Code using Terraform, container orchestration with Kubernetes (managing clusters, nodes, and pods), and building CI/CD pipelines via GitHub Actions. Software & Agentic Engineering (30%): Proficiency with modern scripting languages for building custom AI agents, configuring workflow orchestration, connecting enterprise API integrations, and implementing agentic guardrails. Data Engineering (30%): Practical experience implementing components of data pipelines, including real-time streaming tools (AWS Kinesis, Kafka), data orchestration (dbt, Airflow), and managing vector databases for RAG architectures. Observability & Cost Management: SRE/Platform experience with the practical application of real-time monitoring and cloud cost optimisation using native CSP tools or utilities like Infracost and Karpenter. Domain & Sector Experience Regulated Industries: Experience delivering secure platforms within highly regulated environments, such as iGaming, financial services, or banking, where zero trust security and strict compliance controls are mandatory is highly advantageous. Consulting: Proven experience operating in a client-facing or consulting engineering role, demonstrating strong communication skills, relationship-building capabilities, and adaptability across changing technical contexts. Preferred Knowledge-level & Certifications Associate-level cloud certifications (e.g., AWS Solutions Architect Associate or Azure AZ-104). At least one professional-level certification. Progressing toward or holding multiple professional certifications such as CKA (Certified Kubernetes Administrator), AWS ML Engineer Associate, or FinOps Certified Practitioner. What we'll offer you: We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You'll be part of a supportive and friendly culture, surrounded by smart, curious people who care deeply about what they do. We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed. We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.
Position Overview At CreateFuture, the Senior Cloud Engineer is a vital individual contributor and a technical driving force within our delivery teams. You will work closely with clients and delivery teams to bridge the implementation gap, assisting with architectural designs and transforming them into production systems using autonomous and AI-native platforms. This role requires a modern, hybrid engineering approach, blending infrastructure mastery with software development and data engineering principles. You will act as a technical expert on projects, building trusted relationships with client stakeholders, establishing engineering best practices, and mentoring engineers within the Cloud capability. Furthermore, you will actively contribute to setting engineering standards that influence the whole CreateFuture organisation. Key Responsibilities High-Quality Execution: Proactively contribute to project teams by providing technical guidance, resolving complex technical blockages, supporting migration efforts and writing high-quality, spec-driven code. Platform Engineering: Build and maintain independent, self-service developer platforms that allow engineering teams to spin up compliant, ephemer environments automatically using GitOps workflows. AI & Agentic Infrastructure: Implement the data pipelines, workflow orchestrations, and specialised compute footprints needed to support enterprise AI applications, using technologies like AWS Bedrock and AWS AgentCore. Observability & Reliability: Build robust monitoring and observability pipelines to ensure the health, performance, and security of distributed cloud applications and AI models. FinOps Standards: Embed automated cost-estimation tools into the standard CI/CD deployment pipelines, ensuring auto-shutdown policies and resource optimisation are configured by default. Consultancy & Client Alignment Client Advisory: Collaborate with diverse stakeholders, including tech leads, project managers, and enterprise architects, to ensure successful project delivery according to timelines and budgets. Commercial Growth: Have a growth mindset on projects, recognising potential operational bottlenecks or new requirements, and own communicating these throughout the team for new opportunities. Best Practice Champion: Establish best practice tools and processes within client environments, confidently challenging legacy ways of working where appropriate. Capability & Leadership Support Mentoring: Support and mentor engineers within CreateFuture and our client teams, helping them navigate their learning and technical skill gaps. Community Contribution: Actively engage in the internal engineering communities, helping to build out an internal repository of reference architectures, Infrastructure as Code templates, and technical blogs. Culture and Feedback: Help foster an inclusive, collaborative, and psychologically safe team environment by providing timely, open, and honest technical feedback. Skills & Experience Core Technical Capabilities Candidates must demonstrate a hybrid balance across the three strategic pillars: Software Engineering (30%), Infrastructure (40%), and Data Engineering (30%). Infrastructure & Automation (40%): Lots of experience in writing declarative Infrastructure as Code using Terraform, container orchestration with Kubernetes (managing clusters, nodes, and pods), and building CI/CD pipelines via GitHub Actions. Software & Agentic Engineering (30%): Proficiency with modern scripting languages for building custom AI agents, configuring workflow orchestration, connecting enterprise API integrations, and implementing agentic guardrails. Data Engineering (30%): Practical experience implementing components of data pipelines, including real-time streaming tools (AWS Kinesis, Kafka), data orchestration (dbt, Airflow), and managing vector databases for RAG architectures. Observability & Cost Management: SRE/Platform experience with the practical application of real-time monitoring and cloud cost optimisation using native CSP tools or utilities like Infracost and Karpenter. Domain & Sector Experience Regulated Industries: Experience delivering secure platforms within highly regulated environments, such as iGaming, financial services, or banking, where zero trust security and strict compliance controls are mandatory is highly advantageous. Consulting: Proven experience operating in a client-facing or consulting engineering role, demonstrating strong communication skills, relationship-building capabilities, and adaptability across changing technical contexts. Preferred Knowledge-level & Certifications Associate-level cloud certifications (e.g., AWS Solutions Architect Associate or Azure AZ-104). At least one professional-level certification. Progressing toward or holding multiple professional certifications such as CKA (Certified Kubernetes Administrator), AWS ML Engineer Associate, or FinOps Certified Practitioner. What we'll offer you: We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You'll be part of a supportive and friendly culture, surrounded by smart, curious people who care deeply about what they do. We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed. We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.
21/07/2026
Full time
Position Overview At CreateFuture, the Senior Cloud Engineer is a vital individual contributor and a technical driving force within our delivery teams. You will work closely with clients and delivery teams to bridge the implementation gap, assisting with architectural designs and transforming them into production systems using autonomous and AI-native platforms. This role requires a modern, hybrid engineering approach, blending infrastructure mastery with software development and data engineering principles. You will act as a technical expert on projects, building trusted relationships with client stakeholders, establishing engineering best practices, and mentoring engineers within the Cloud capability. Furthermore, you will actively contribute to setting engineering standards that influence the whole CreateFuture organisation. Key Responsibilities High-Quality Execution: Proactively contribute to project teams by providing technical guidance, resolving complex technical blockages, supporting migration efforts and writing high-quality, spec-driven code. Platform Engineering: Build and maintain independent, self-service developer platforms that allow engineering teams to spin up compliant, ephemer environments automatically using GitOps workflows. AI & Agentic Infrastructure: Implement the data pipelines, workflow orchestrations, and specialised compute footprints needed to support enterprise AI applications, using technologies like AWS Bedrock and AWS AgentCore. Observability & Reliability: Build robust monitoring and observability pipelines to ensure the health, performance, and security of distributed cloud applications and AI models. FinOps Standards: Embed automated cost-estimation tools into the standard CI/CD deployment pipelines, ensuring auto-shutdown policies and resource optimisation are configured by default. Consultancy & Client Alignment Client Advisory: Collaborate with diverse stakeholders, including tech leads, project managers, and enterprise architects, to ensure successful project delivery according to timelines and budgets. Commercial Growth: Have a growth mindset on projects, recognising potential operational bottlenecks or new requirements, and own communicating these throughout the team for new opportunities. Best Practice Champion: Establish best practice tools and processes within client environments, confidently challenging legacy ways of working where appropriate. Capability & Leadership Support Mentoring: Support and mentor engineers within CreateFuture and our client teams, helping them navigate their learning and technical skill gaps. Community Contribution: Actively engage in the internal engineering communities, helping to build out an internal repository of reference architectures, Infrastructure as Code templates, and technical blogs. Culture and Feedback: Help foster an inclusive, collaborative, and psychologically safe team environment by providing timely, open, and honest technical feedback. Skills & Experience Core Technical Capabilities Candidates must demonstrate a hybrid balance across the three strategic pillars: Software Engineering (30%), Infrastructure (40%), and Data Engineering (30%). Infrastructure & Automation (40%): Lots of experience in writing declarative Infrastructure as Code using Terraform, container orchestration with Kubernetes (managing clusters, nodes, and pods), and building CI/CD pipelines via GitHub Actions. Software & Agentic Engineering (30%): Proficiency with modern scripting languages for building custom AI agents, configuring workflow orchestration, connecting enterprise API integrations, and implementing agentic guardrails. Data Engineering (30%): Practical experience implementing components of data pipelines, including real-time streaming tools (AWS Kinesis, Kafka), data orchestration (dbt, Airflow), and managing vector databases for RAG architectures. Observability & Cost Management: SRE/Platform experience with the practical application of real-time monitoring and cloud cost optimisation using native CSP tools or utilities like Infracost and Karpenter. Domain & Sector Experience Regulated Industries: Experience delivering secure platforms within highly regulated environments, such as iGaming, financial services, or banking, where zero trust security and strict compliance controls are mandatory is highly advantageous. Consulting: Proven experience operating in a client-facing or consulting engineering role, demonstrating strong communication skills, relationship-building capabilities, and adaptability across changing technical contexts. Preferred Knowledge-level & Certifications Associate-level cloud certifications (e.g., AWS Solutions Architect Associate or Azure AZ-104). At least one professional-level certification. Progressing toward or holding multiple professional certifications such as CKA (Certified Kubernetes Administrator), AWS ML Engineer Associate, or FinOps Certified Practitioner. What we'll offer you: We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You'll be part of a supportive and friendly culture, surrounded by smart, curious people who care deeply about what they do. We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed. We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.
Position Overview At CreateFuture, the Senior Cloud Engineer is a vital individual contributor and a technical driving force within our delivery teams. You will work closely with clients and delivery teams to bridge the implementation gap, assisting with architectural designs and transforming them into production systems using autonomous and AI-native platforms. This role requires a modern, hybrid engineering approach, blending infrastructure mastery with software development and data engineering principles. You will act as a technical expert on projects, building trusted relationships with client stakeholders, establishing engineering best practices, and mentoring engineers within the Cloud capability. Furthermore, you will actively contribute to setting engineering standards that influence the whole CreateFuture organisation. Key Responsibilities High-Quality Execution: Proactively contribute to project teams by providing technical guidance, resolving complex technical blockages, supporting migration efforts and writing high-quality, spec-driven code. Platform Engineering: Build and maintain independent, self-service developer platforms that allow engineering teams to spin up compliant, ephemer environments automatically using GitOps workflows. AI & Agentic Infrastructure: Implement the data pipelines, workflow orchestrations, and specialised compute footprints needed to support enterprise AI applications, using technologies like AWS Bedrock and AWS AgentCore. Observability & Reliability: Build robust monitoring and observability pipelines to ensure the health, performance, and security of distributed cloud applications and AI models. FinOps Standards: Embed automated cost-estimation tools into the standard CI/CD deployment pipelines, ensuring auto-shutdown policies and resource optimisation are configured by default. Consultancy & Client Alignment Client Advisory: Collaborate with diverse stakeholders, including tech leads, project managers, and enterprise architects, to ensure successful project delivery according to timelines and budgets. Commercial Growth: Have a growth mindset on projects, recognising potential operational bottlenecks or new requirements, and own communicating these throughout the team for new opportunities. Best Practice Champion: Establish best practice tools and processes within client environments, confidently challenging legacy ways of working where appropriate. Capability & Leadership Support Mentoring: Support and mentor engineers within CreateFuture and our client teams, helping them navigate their learning and technical skill gaps. Community Contribution: Actively engage in the internal engineering communities, helping to build out an internal repository of reference architectures, Infrastructure as Code templates, and technical blogs. Culture and Feedback: Help foster an inclusive, collaborative, and psychologically safe team environment by providing timely, open, and honest technical feedback. Skills & Experience Core Technical Capabilities Candidates must demonstrate a hybrid balance across the three strategic pillars: Software Engineering (30%), Infrastructure (40%), and Data Engineering (30%). Infrastructure & Automation (40%): Lots of experience in writing declarative Infrastructure as Code using Terraform, container orchestration with Kubernetes (managing clusters, nodes, and pods), and building CI/CD pipelines via GitHub Actions. Software & Agentic Engineering (30%): Proficiency with modern scripting languages for building custom AI agents, configuring workflow orchestration, connecting enterprise API integrations, and implementing agentic guardrails. Data Engineering (30%): Practical experience implementing components of data pipelines, including real-time streaming tools (AWS Kinesis, Kafka), data orchestration (dbt, Airflow), and managing vector databases for RAG architectures. Observability & Cost Management: SRE/Platform experience with the practical application of real-time monitoring and cloud cost optimisation using native CSP tools or utilities like Infracost and Karpenter. Domain & Sector Experience Regulated Industries: Experience delivering secure platforms within highly regulated environments, such as iGaming, financial services, or banking, where zero trust security and strict compliance controls are mandatory is highly advantageous. Consulting: Proven experience operating in a client-facing or consulting engineering role, demonstrating strong communication skills, relationship-building capabilities, and adaptability across changing technical contexts. Preferred Knowledge-level & Certifications Associate-level cloud certifications (e.g., AWS Solutions Architect Associate or Azure AZ-104). At least one professional-level certification. Progressing toward or holding multiple professional certifications such as CKA (Certified Kubernetes Administrator), AWS ML Engineer Associate, or FinOps Certified Practitioner. What we'll offer you: We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You'll be part of a supportive and friendly culture, surrounded by smart, curious people who care deeply about what they do. We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed. We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.
21/07/2026
Full time
Position Overview At CreateFuture, the Senior Cloud Engineer is a vital individual contributor and a technical driving force within our delivery teams. You will work closely with clients and delivery teams to bridge the implementation gap, assisting with architectural designs and transforming them into production systems using autonomous and AI-native platforms. This role requires a modern, hybrid engineering approach, blending infrastructure mastery with software development and data engineering principles. You will act as a technical expert on projects, building trusted relationships with client stakeholders, establishing engineering best practices, and mentoring engineers within the Cloud capability. Furthermore, you will actively contribute to setting engineering standards that influence the whole CreateFuture organisation. Key Responsibilities High-Quality Execution: Proactively contribute to project teams by providing technical guidance, resolving complex technical blockages, supporting migration efforts and writing high-quality, spec-driven code. Platform Engineering: Build and maintain independent, self-service developer platforms that allow engineering teams to spin up compliant, ephemer environments automatically using GitOps workflows. AI & Agentic Infrastructure: Implement the data pipelines, workflow orchestrations, and specialised compute footprints needed to support enterprise AI applications, using technologies like AWS Bedrock and AWS AgentCore. Observability & Reliability: Build robust monitoring and observability pipelines to ensure the health, performance, and security of distributed cloud applications and AI models. FinOps Standards: Embed automated cost-estimation tools into the standard CI/CD deployment pipelines, ensuring auto-shutdown policies and resource optimisation are configured by default. Consultancy & Client Alignment Client Advisory: Collaborate with diverse stakeholders, including tech leads, project managers, and enterprise architects, to ensure successful project delivery according to timelines and budgets. Commercial Growth: Have a growth mindset on projects, recognising potential operational bottlenecks or new requirements, and own communicating these throughout the team for new opportunities. Best Practice Champion: Establish best practice tools and processes within client environments, confidently challenging legacy ways of working where appropriate. Capability & Leadership Support Mentoring: Support and mentor engineers within CreateFuture and our client teams, helping them navigate their learning and technical skill gaps. Community Contribution: Actively engage in the internal engineering communities, helping to build out an internal repository of reference architectures, Infrastructure as Code templates, and technical blogs. Culture and Feedback: Help foster an inclusive, collaborative, and psychologically safe team environment by providing timely, open, and honest technical feedback. Skills & Experience Core Technical Capabilities Candidates must demonstrate a hybrid balance across the three strategic pillars: Software Engineering (30%), Infrastructure (40%), and Data Engineering (30%). Infrastructure & Automation (40%): Lots of experience in writing declarative Infrastructure as Code using Terraform, container orchestration with Kubernetes (managing clusters, nodes, and pods), and building CI/CD pipelines via GitHub Actions. Software & Agentic Engineering (30%): Proficiency with modern scripting languages for building custom AI agents, configuring workflow orchestration, connecting enterprise API integrations, and implementing agentic guardrails. Data Engineering (30%): Practical experience implementing components of data pipelines, including real-time streaming tools (AWS Kinesis, Kafka), data orchestration (dbt, Airflow), and managing vector databases for RAG architectures. Observability & Cost Management: SRE/Platform experience with the practical application of real-time monitoring and cloud cost optimisation using native CSP tools or utilities like Infracost and Karpenter. Domain & Sector Experience Regulated Industries: Experience delivering secure platforms within highly regulated environments, such as iGaming, financial services, or banking, where zero trust security and strict compliance controls are mandatory is highly advantageous. Consulting: Proven experience operating in a client-facing or consulting engineering role, demonstrating strong communication skills, relationship-building capabilities, and adaptability across changing technical contexts. Preferred Knowledge-level & Certifications Associate-level cloud certifications (e.g., AWS Solutions Architect Associate or Azure AZ-104). At least one professional-level certification. Progressing toward or holding multiple professional certifications such as CKA (Certified Kubernetes Administrator), AWS ML Engineer Associate, or FinOps Certified Practitioner. What we'll offer you: We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You'll be part of a supportive and friendly culture, surrounded by smart, curious people who care deeply about what they do. We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed. We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.
Role Purpose The Site Reliability Engineer will work closely with Application, Infrastructure, and Network Engineering teams to ensure the reliability, scalability, and performance of FNZ platforms. This role focuses on deploying, integrating, and providing ongoing operational support for mission-critical systems, leveraging modern automation and cloud-native practices. Key Responsibilities Maintain high availability and performance of FNZ platforms. Implement monitoring, alerting, and observability solutions to proactively detect and resolve issues. Collaborate with engineering teams to design and implement robust deployment pipelines. Ensure smooth integration of applications with infrastructure and network components. Use Terraform for provisioning and managing infrastructure across environments. Operate and optimize workloads on-prem and public cloud. Manage and troubleshoot application delivery networks, load balancing, and traffic routing. Configure and support F5 Distributed Cloud or similar CDN/ADC technologies. Participate in on-call rotations, perform root-cause analysis, and implement preventive measures. Work cross-functionally with Application, Infrastructure, and Network Engineering teams to deliver reliable services. Required Skills & Experience Kubernetes (K8s): Deep understanding of container orchestration and cluster management. Terraform: Strong experience in Infrastructure as Code for cloud and on-prem environments. Public Cloud: Hands-on experience with AWS, Azure, or GCP. F5 Distributed Cloud or Similar: Knowledge of CDN/ADC platforms and their integration. Networking Fundamentals: Expertise in application delivery networks, load balancing, traffic routing, and troubleshooting. Observability Tools: Familiarity with Splunk, NewRelic, or similar. Scripting & Automation: Proficiency in Terraform, Bash, or similar languages. Desirable Skills Experience with CI/CD pipelines and GitOps workflows. Knowledge of SRE principles. Familiarity with security best practices. Key Attributes Strong problem-solving and troubleshooting skills. Ability to work collaboratively across multiple teams. Passion for automation and reducing operational toil. Reporting Line Reports to: Head of Platform Operations/Application Engineering. Works closely with Application Engineering, Infrastructure Engineering, and Network Engineering teams.
19/07/2026
Full time
Role Purpose The Site Reliability Engineer will work closely with Application, Infrastructure, and Network Engineering teams to ensure the reliability, scalability, and performance of FNZ platforms. This role focuses on deploying, integrating, and providing ongoing operational support for mission-critical systems, leveraging modern automation and cloud-native practices. Key Responsibilities Maintain high availability and performance of FNZ platforms. Implement monitoring, alerting, and observability solutions to proactively detect and resolve issues. Collaborate with engineering teams to design and implement robust deployment pipelines. Ensure smooth integration of applications with infrastructure and network components. Use Terraform for provisioning and managing infrastructure across environments. Operate and optimize workloads on-prem and public cloud. Manage and troubleshoot application delivery networks, load balancing, and traffic routing. Configure and support F5 Distributed Cloud or similar CDN/ADC technologies. Participate in on-call rotations, perform root-cause analysis, and implement preventive measures. Work cross-functionally with Application, Infrastructure, and Network Engineering teams to deliver reliable services. Required Skills & Experience Kubernetes (K8s): Deep understanding of container orchestration and cluster management. Terraform: Strong experience in Infrastructure as Code for cloud and on-prem environments. Public Cloud: Hands-on experience with AWS, Azure, or GCP. F5 Distributed Cloud or Similar: Knowledge of CDN/ADC platforms and their integration. Networking Fundamentals: Expertise in application delivery networks, load balancing, traffic routing, and troubleshooting. Observability Tools: Familiarity with Splunk, NewRelic, or similar. Scripting & Automation: Proficiency in Terraform, Bash, or similar languages. Desirable Skills Experience with CI/CD pipelines and GitOps workflows. Knowledge of SRE principles. Familiarity with security best practices. Key Attributes Strong problem-solving and troubleshooting skills. Ability to work collaboratively across multiple teams. Passion for automation and reducing operational toil. Reporting Line Reports to: Head of Platform Operations/Application Engineering. Works closely with Application Engineering, Infrastructure Engineering, and Network Engineering teams.
Select how often (in days) to receive an alert: Create Alert Location: Hybrid, with a minimum of 20% in the London office per month About Us: We're Nominet - a world-leading domain name registry operating at the heart of the UK internet. While we're best known for running .UK domains, our DNS expertise also underpins critical internet infrastructure that government services, including the NHS, rely on. As a public benefit company, our work has a positive impact on society. We've donated millions to projects that use technology to improve people's lives and have committed to delivering £60m worth of support over the next three years. The Role: We're hiring a Site Reliability Engineer to join our Reliability Engineering team. This team builds and runs the secure platforms behind Nominet's registry and DNS services. The role focuses on AWS, Kubernetes, infrastructure as code, GitOps, CI/CD, and self-service tooling. You'll help design, deploy, and operate scalable cloud infrastructure. You'll also improve how developers use that infrastructure, reducing toil and making delivery safer, simpler, and more repeatable. This is a role for someone who can build systems, not just operate them. You'll need sound technical judgment, clear communication, and the confidence to challenge decisions when reliability, security, or simplicity is at risk. What You'll Be Doing: Help build and maintain internal tooling in Go Design, deploy, and manage Kubernetes clusters that support scalable and resilient application infrastructure Build and improve self-service infrastructure tools that help developers provision and manage resources safely Implement and maintain CI/CD pipelines that make application delivery clearer and more reliable Use Terraform to manage infrastructure as code and support automated provisioning Apply GitOps principles to infrastructure deployment and configuration Work with development teams to integrate services into production AWS environments Monitor, troubleshoot, and improve the performance of infrastructure components Help keep platforms secure, scalable, stable, and aligned to compliance requirements Build scripts and tools in a modern programming language to automate routine operations. Bash scripting is useful as a secondary skill Maintain clear documentation for infrastructure setups, procedures, and architectural decisions Work closely with security teams to apply practical security controls and good engineering practice Assess cloud, container, and GitOps tools where they can reduce toil, improve reliability, or make the developer experience simpler About You: You'll bring practical experience in SRE, platform engineering, DevOps, or cloud engineering. You do not need to have followed one fixed career path, but you should be able to show evidence of running and improving production infrastructure. You'll likely have experience with: Operating production systems on AWS Experience with Go or another language (We use Go and will need someone who wants to use this language) Building or running Kubernetes platforms in production Creating or improving CI/CD pipelines, ideally with GitLab or a similar tool Applying GitOps workflows, ideally with ArgoCD or similar tooling Automating infrastructure tasks using a modern programming language Working with developers, security teams, and other engineers to solve platform problems Certifications such as AWS Certified Solutions Architect, AWS Certified DevOps Engineer, CKA, or CKAD may be useful, but they are not a substitute for practical experience What To Expect Next: 1st stage: Introduction call with a member of the TA team (30 mins) 2nd stage: Hiring Manager Interview (60 mins via Teams) 3rd stage: On-site interview, including technical tasks (90-120 mins) What We Offer: Hybrid & Flexible Working Early Finish Friday - Working week of 34 hours with full-time pay (Finish at midday on Friday) 30 days of annual leave plus bank holidays, with the ability to purchase an additional 5 days Private Medical Insurance + Employee Assistance Programme Pension Scheme (Matched to 7%) Annual Bonus Scheme Family Leave (Enhanced) Electric vehicle scheme with on-site charging points Rewards platform with access to discounts at hundreds of shops, restaurants etc. Flexible Benefits Diversity Statement: We're passionate about creating a workplace where every individual is valued, respected, and empowered. Somewhere we can benefit from all forms of diversity and discover the true value in our differences. If there are any adjustments we could make to the recruitment and selection process to support you, please let us know Security Statement Nominet is committed to the safeguarding and welfare of the internet and expects all employees and volunteers to share this commitment by participating in the relevant security and screening processes. All roles working for Nominet will be subject to a Baseline Personnel Security Standard (BPSS) check. Some roles due to the nature of their work, will require additional security clearance.
19/07/2026
Full time
Select how often (in days) to receive an alert: Create Alert Location: Hybrid, with a minimum of 20% in the London office per month About Us: We're Nominet - a world-leading domain name registry operating at the heart of the UK internet. While we're best known for running .UK domains, our DNS expertise also underpins critical internet infrastructure that government services, including the NHS, rely on. As a public benefit company, our work has a positive impact on society. We've donated millions to projects that use technology to improve people's lives and have committed to delivering £60m worth of support over the next three years. The Role: We're hiring a Site Reliability Engineer to join our Reliability Engineering team. This team builds and runs the secure platforms behind Nominet's registry and DNS services. The role focuses on AWS, Kubernetes, infrastructure as code, GitOps, CI/CD, and self-service tooling. You'll help design, deploy, and operate scalable cloud infrastructure. You'll also improve how developers use that infrastructure, reducing toil and making delivery safer, simpler, and more repeatable. This is a role for someone who can build systems, not just operate them. You'll need sound technical judgment, clear communication, and the confidence to challenge decisions when reliability, security, or simplicity is at risk. What You'll Be Doing: Help build and maintain internal tooling in Go Design, deploy, and manage Kubernetes clusters that support scalable and resilient application infrastructure Build and improve self-service infrastructure tools that help developers provision and manage resources safely Implement and maintain CI/CD pipelines that make application delivery clearer and more reliable Use Terraform to manage infrastructure as code and support automated provisioning Apply GitOps principles to infrastructure deployment and configuration Work with development teams to integrate services into production AWS environments Monitor, troubleshoot, and improve the performance of infrastructure components Help keep platforms secure, scalable, stable, and aligned to compliance requirements Build scripts and tools in a modern programming language to automate routine operations. Bash scripting is useful as a secondary skill Maintain clear documentation for infrastructure setups, procedures, and architectural decisions Work closely with security teams to apply practical security controls and good engineering practice Assess cloud, container, and GitOps tools where they can reduce toil, improve reliability, or make the developer experience simpler About You: You'll bring practical experience in SRE, platform engineering, DevOps, or cloud engineering. You do not need to have followed one fixed career path, but you should be able to show evidence of running and improving production infrastructure. You'll likely have experience with: Operating production systems on AWS Experience with Go or another language (We use Go and will need someone who wants to use this language) Building or running Kubernetes platforms in production Creating or improving CI/CD pipelines, ideally with GitLab or a similar tool Applying GitOps workflows, ideally with ArgoCD or similar tooling Automating infrastructure tasks using a modern programming language Working with developers, security teams, and other engineers to solve platform problems Certifications such as AWS Certified Solutions Architect, AWS Certified DevOps Engineer, CKA, or CKAD may be useful, but they are not a substitute for practical experience What To Expect Next: 1st stage: Introduction call with a member of the TA team (30 mins) 2nd stage: Hiring Manager Interview (60 mins via Teams) 3rd stage: On-site interview, including technical tasks (90-120 mins) What We Offer: Hybrid & Flexible Working Early Finish Friday - Working week of 34 hours with full-time pay (Finish at midday on Friday) 30 days of annual leave plus bank holidays, with the ability to purchase an additional 5 days Private Medical Insurance + Employee Assistance Programme Pension Scheme (Matched to 7%) Annual Bonus Scheme Family Leave (Enhanced) Electric vehicle scheme with on-site charging points Rewards platform with access to discounts at hundreds of shops, restaurants etc. Flexible Benefits Diversity Statement: We're passionate about creating a workplace where every individual is valued, respected, and empowered. Somewhere we can benefit from all forms of diversity and discover the true value in our differences. If there are any adjustments we could make to the recruitment and selection process to support you, please let us know Security Statement Nominet is committed to the safeguarding and welfare of the internet and expects all employees and volunteers to share this commitment by participating in the relevant security and screening processes. All roles working for Nominet will be subject to a Baseline Personnel Security Standard (BPSS) check. Some roles due to the nature of their work, will require additional security clearance.
Responsibilities Design, build and operate scalable, resilient GKE environments Engineer multi-tenant Kubernetes clusters with strong workload isolation and platform guardrails Support shared and dedicated cluster patterns, including tenant onboarding Improve platform performance under production conditions (e.g. scaling, storage, node pressure) Build automation-first infrastructure using Terraform, CI/CD and GitOps Simplify cluster lifecycle management (provisioning, upgrades, add-ons) Develop self-service platform capabilities to improve developer experience Apply Site Reliability Engineering (SRE) practices to platform operations Support incident response, monitoring, observability and continuous improvement Diagnose issues across performance, scaling, storage and automation Contribute to a 24x7 on-call rotation Implement policy-as-code controls (e.g. OPA Gatekeeper, RBAC, workload identity) Support audit, compliance and risk mitigation activities Ensure platforms are secure, supportable and aligned to control frameworks Work with service mesh and ingress/egress patterns (e.g. Istio, Anthos, Cloud Service Mesh) Support cloud networking (VPCs, DNS, NAT, VPN, routing, connectivity) Integrate shared platform services (cert-manager, observability, cost tooling) Requirements Strong experience in Platform Engineering, DevOps or SRE Proven delivery of production Kubernetes platforms, ideally GKE Experience with multi-tenant platform environments (shared clusters, isolation, scaling) Deep understanding of Kubernetes internals (scheduling, storage, node lifecycle, upgrades) Strong knowledge of Google Cloud Platform (GCP), including: GKE IAM / Workload Identity Networking (VPC, DNS, NAT, ingress/egress) Storage patterns Experience with Infrastructure as Code (Terraform) using modular design Strong experience with CI/CD pipelines and GitOps workflows Coding/scripting (Python, Go or Bash) Strong troubleshooting and problem-solving skills Ability to own and deliver complex engineering outcomes
19/07/2026
Full time
Responsibilities Design, build and operate scalable, resilient GKE environments Engineer multi-tenant Kubernetes clusters with strong workload isolation and platform guardrails Support shared and dedicated cluster patterns, including tenant onboarding Improve platform performance under production conditions (e.g. scaling, storage, node pressure) Build automation-first infrastructure using Terraform, CI/CD and GitOps Simplify cluster lifecycle management (provisioning, upgrades, add-ons) Develop self-service platform capabilities to improve developer experience Apply Site Reliability Engineering (SRE) practices to platform operations Support incident response, monitoring, observability and continuous improvement Diagnose issues across performance, scaling, storage and automation Contribute to a 24x7 on-call rotation Implement policy-as-code controls (e.g. OPA Gatekeeper, RBAC, workload identity) Support audit, compliance and risk mitigation activities Ensure platforms are secure, supportable and aligned to control frameworks Work with service mesh and ingress/egress patterns (e.g. Istio, Anthos, Cloud Service Mesh) Support cloud networking (VPCs, DNS, NAT, VPN, routing, connectivity) Integrate shared platform services (cert-manager, observability, cost tooling) Requirements Strong experience in Platform Engineering, DevOps or SRE Proven delivery of production Kubernetes platforms, ideally GKE Experience with multi-tenant platform environments (shared clusters, isolation, scaling) Deep understanding of Kubernetes internals (scheduling, storage, node lifecycle, upgrades) Strong knowledge of Google Cloud Platform (GCP), including: GKE IAM / Workload Identity Networking (VPC, DNS, NAT, ingress/egress) Storage patterns Experience with Infrastructure as Code (Terraform) using modular design Strong experience with CI/CD pipelines and GitOps workflows Coding/scripting (Python, Go or Bash) Strong troubleshooting and problem-solving skills Ability to own and deliver complex engineering outcomes
Please note that these positions are based in London, Berlin or Paris - relocation support is provided if required. THE BEST WORK OF YOUR CAREER Trade Republic is the largest savings platform in Europe - we operate in 18 countries, serving million customers who trust us with over €150B in assets. But we're striving for more. We have a bold mission to empower everyone to build wealth with easy, safe, and free access to financial systems. You will have the opportunity to grow your career by collaborating with a team of outstanding talents and state of the art technology to build a lasting, positive future for millions. ABOUT PLATFORM ENGINEERING Platform Engineering is the backbone of Trade Republic's engineering velocity. Our mission is to build scalable platforms for a Europe scale bank - serving internal engineers, and building in house control planes for managing the bank's infrastructure. We're a 50 person Platform team focused on one thing: enabling product engineers to move fast and operate autonomously by default. We build self service platforms, golden paths, and opinionated tooling so that over 400 engineers can ship with confidence. From Kubernetes fleet management and CI/CD to an internal Developer Hub built on Backstage, our work underpins every trade, savings plan, and card payment that flows through the platform. THE KUBERNETES PLATFORM JOURNEY Trade Republic runs a growing fleet of Kubernetes clusters across multiple AWS accounts and environments, underpinning every service that millions of customers depend on. Managing that fleet at the pace the business demands - with the reliability and compliance bar a regulated financial institution requires - made it clear that treating clusters as manually managed infrastructure was no longer an option. We changed that. We built a hub and spoke fleet management platform - ClusterAPI for cluster lifecycle, Crossplane for the infrastructure control plane, Sveltos for addon orchestration, all driven by GitOps from a single source of truth. A new cluster is now a single resource definition. Addons are versioned packages, deployed automatically to every cluster in the fleet. Operational complexity no longer grows with the number of clusters. The foundation is in place. Now we need to take it further: hardening the platform, expanding to multi region, pushing availability targets to meet mission critical service requirements, and making the fleet truly self healing by default. You will lead that next chapter. WHAT YOU'LL BE DOING As a Kubernetes Senior Engineer, you will build and operate the infrastructure layer that every engineer at Trade Republic depends on to deploy and scale their services with confidence. Build for availability: Implement and improve the cluster topology and isolation model that gets mission critical services to their availability targets - dedicated clusters per criticality tier, strict failure boundaries between platform and product workloads, and disaster recovery procedures you help test and prove. Work across the networking fabric: Build and maintain the multi account network the fleet runs on - VPC design, private endpoints, and cross account connectivity - and contribute to the path toward cross region and multi cloud. Make sure services communicate reliably and securely wherever they run. Push toward zero touch fleet operations: Build the automation that turns cluster and addon management into self healing workflows, freeing the team from repetitive operational toil so our time goes toward raising the platform's reliability and capabilities rather than keeping it running. Strengthen the compute contract with product teams: Implement the resource policies, cluster profiles, and workload isolation boundaries that define how workloads run on the platform, and build the guardrails that prevent noisy neighbours to compromise resiliency. Own what you build end to end: Participate in the on call rotation, taking operational ownership of the systems you build and run. Contribute to the platform's direction: Bring strong technical input to the Kubernetes platform's evolution, take features from kickoff to delivery, and weigh infrastructure decisions against reliability, compliance, and cost. WHAT WE'RE LOOKING FOR 5+ years of experience in infrastructure, platform engineering, or a related SRE/systems discipline. Deep, hands on Kubernetes expertise: cluster architecture, scaling, security, and HA design in production environments. Strong networking fundamentals: CNI, ingress, DNS, service mesh, and cross region/cross cloud connectivity. Experience operating multi region and/or multi cloud environments in production. Experience automating the cluster lifecycle - provisioning, upgrades, policy, and governance - with self service enablement for developers. Solid backend engineering background - Go is preferred but any strong language is welcome. A pragmatic approach to applying infrastructure best practices in a way engineering teams actually adopt. Ability to contribute to architectural decisions and clearly communicate trade offs. The ability to work in a flexible hybrid setup, with 2-3 days a week in the office. WHY YOU SHOULD APPLY NOW Our culture rewards ownership, excellence, and high energy. We care deeply about outcomes and hold each other accountable - we're here to win and fix one of the largest challenges Europeans face - closing the pension gap and democratising wealth. If this gets you fired up, reach out! We believe it's our team's varied identities and backgrounds that make us sharper and stronger. We're committed to creating an environment where everyone feels respected and has equal opportunity to thrive in their careers. For any questions on DEI during the interview process, reach out to your recruitment partner.
18/07/2026
Full time
Please note that these positions are based in London, Berlin or Paris - relocation support is provided if required. THE BEST WORK OF YOUR CAREER Trade Republic is the largest savings platform in Europe - we operate in 18 countries, serving million customers who trust us with over €150B in assets. But we're striving for more. We have a bold mission to empower everyone to build wealth with easy, safe, and free access to financial systems. You will have the opportunity to grow your career by collaborating with a team of outstanding talents and state of the art technology to build a lasting, positive future for millions. ABOUT PLATFORM ENGINEERING Platform Engineering is the backbone of Trade Republic's engineering velocity. Our mission is to build scalable platforms for a Europe scale bank - serving internal engineers, and building in house control planes for managing the bank's infrastructure. We're a 50 person Platform team focused on one thing: enabling product engineers to move fast and operate autonomously by default. We build self service platforms, golden paths, and opinionated tooling so that over 400 engineers can ship with confidence. From Kubernetes fleet management and CI/CD to an internal Developer Hub built on Backstage, our work underpins every trade, savings plan, and card payment that flows through the platform. THE KUBERNETES PLATFORM JOURNEY Trade Republic runs a growing fleet of Kubernetes clusters across multiple AWS accounts and environments, underpinning every service that millions of customers depend on. Managing that fleet at the pace the business demands - with the reliability and compliance bar a regulated financial institution requires - made it clear that treating clusters as manually managed infrastructure was no longer an option. We changed that. We built a hub and spoke fleet management platform - ClusterAPI for cluster lifecycle, Crossplane for the infrastructure control plane, Sveltos for addon orchestration, all driven by GitOps from a single source of truth. A new cluster is now a single resource definition. Addons are versioned packages, deployed automatically to every cluster in the fleet. Operational complexity no longer grows with the number of clusters. The foundation is in place. Now we need to take it further: hardening the platform, expanding to multi region, pushing availability targets to meet mission critical service requirements, and making the fleet truly self healing by default. You will lead that next chapter. WHAT YOU'LL BE DOING As a Kubernetes Senior Engineer, you will build and operate the infrastructure layer that every engineer at Trade Republic depends on to deploy and scale their services with confidence. Build for availability: Implement and improve the cluster topology and isolation model that gets mission critical services to their availability targets - dedicated clusters per criticality tier, strict failure boundaries between platform and product workloads, and disaster recovery procedures you help test and prove. Work across the networking fabric: Build and maintain the multi account network the fleet runs on - VPC design, private endpoints, and cross account connectivity - and contribute to the path toward cross region and multi cloud. Make sure services communicate reliably and securely wherever they run. Push toward zero touch fleet operations: Build the automation that turns cluster and addon management into self healing workflows, freeing the team from repetitive operational toil so our time goes toward raising the platform's reliability and capabilities rather than keeping it running. Strengthen the compute contract with product teams: Implement the resource policies, cluster profiles, and workload isolation boundaries that define how workloads run on the platform, and build the guardrails that prevent noisy neighbours to compromise resiliency. Own what you build end to end: Participate in the on call rotation, taking operational ownership of the systems you build and run. Contribute to the platform's direction: Bring strong technical input to the Kubernetes platform's evolution, take features from kickoff to delivery, and weigh infrastructure decisions against reliability, compliance, and cost. WHAT WE'RE LOOKING FOR 5+ years of experience in infrastructure, platform engineering, or a related SRE/systems discipline. Deep, hands on Kubernetes expertise: cluster architecture, scaling, security, and HA design in production environments. Strong networking fundamentals: CNI, ingress, DNS, service mesh, and cross region/cross cloud connectivity. Experience operating multi region and/or multi cloud environments in production. Experience automating the cluster lifecycle - provisioning, upgrades, policy, and governance - with self service enablement for developers. Solid backend engineering background - Go is preferred but any strong language is welcome. A pragmatic approach to applying infrastructure best practices in a way engineering teams actually adopt. Ability to contribute to architectural decisions and clearly communicate trade offs. The ability to work in a flexible hybrid setup, with 2-3 days a week in the office. WHY YOU SHOULD APPLY NOW Our culture rewards ownership, excellence, and high energy. We care deeply about outcomes and hold each other accountable - we're here to win and fix one of the largest challenges Europeans face - closing the pension gap and democratising wealth. If this gets you fired up, reach out! We believe it's our team's varied identities and backgrounds that make us sharper and stronger. We're committed to creating an environment where everyone feels respected and has equal opportunity to thrive in their careers. For any questions on DEI during the interview process, reach out to your recruitment partner.
Senior Platform Engineer Salary: £100,000 - £110,000 base salary + benefits Build the platforms behind critical digital services Are you a Senior or Lead Platform Engineer who thrives on solving complex infrastructure challenges, building production-grade platforms, and shaping the engineering practices that enable teams to deliver at scale? We are looking for experienced Platform Engineers to help design, build and continuously improve critical digital services used across the Nation. You'll work on platforms that must remain secure, resilient and observable under significant demand, helping to modernise essential services that make a real difference. You'll join a highly collaborative engineering community where knowledge sharing, continuous improvement and technical excellence are at the heart of everything we do. The Opportunity As a Senior Platform Engineer, you'll play a key role in designing, building and evolving modern cloud platforms that enable engineering teams to deliver safely, rapidly and reliably. This is a hands-on technical leadership role where you'll combine deep engineering expertise with strategic influence. You'll work closely with delivery teams, architects and client stakeholders to define platform strategy, establish engineering standards and build reusable capabilities that improve the developer experience. You'll remain close to the technology, actively contributing to architecture, code, automation and operational improvements while mentoring and supporting other engineers. What You'll Be Doing You'll help design, build and operate the platforms that underpin mission-critical services, including: Designing secure, scalable multi-cloud landing zones, primarily within AWS. Building GitOps-driven platforms and Internal Developer Platforms (IDPs) that provide true self-service capabilities for engineering teams. Modernising legacy environments into cloud-native architectures. Creating unified observability solutions using OpenTelemetry, Prometheus, Grafana and modern APM tooling. Driving improvements in reliability, operability and platform performance through Site Reliability Engineering (SRE) practices. Implementing secure CI/CD pipelines with a strong focus on DevSecOps and software supply-chain integrity. Delivering event-driven automation and infrastructure capabilities. Driving FinOps initiatives and optimising cloud consumption across large-scale estates. Applying platform-as-a-product principles to create reusable, scalable engineering capabilities. Technical Leadership & Influence As a senior member of the engineering team, you will: Act as a senior technical authority, influencing platform strategy, architecture and governance. Lead technical workshops, architecture discussions and design reviews with stakeholders and delivery teams. Define and evolve engineering standards covering security, reliability, observability, CI/CD and operational excellence. Establish Service Level Objectives (SLOs), improve reliability and facilitate incident reviews and continuous improvement activities. Champion platform engineering best practices and help teams adopt them successfully. Coach and mentor engineers through pairing, technical leadership and knowledge sharing. Drive platform-as-a-product thinking, establishing golden paths and self-service capabilities that improve developer experience. Remain hands-on by contributing to designs, code reviews, automation and complex technical problem solving. What We're Looking For Essential Experience Significant experience designing, building and operating cloud-native platforms in enterprise or large-scale environments. Strong experience with AWS, including multi-account architectures and secure landing zones. Experience building and operating Kubernetes platforms, particularly EKS (AKS experience also beneficial). Expertise in Infrastructure as Code using tools such as Terraform. Strong understanding of CI/CD, DevSecOps and modern software delivery practices. Experience implementing observability solutions using tools such as OpenTelemetry, Prometheus and Grafana. Strong knowledge of cloud security, identity, networking and platform reliability. Experience leading technical discussions and influencing engineering direction across multiple teams. Proven ability to mentor engineers and provide technical leadership within multidisciplinary teams. Desirable Experience Experience with Azure and/or Google Cloud Platform. Experience building Internal Developer Platforms and self-service engineering capabilities. Knowledge of GitOps tooling such as Argo CD or Flux. Experience applying Site Reliability Engineering (SRE) practices. Experience driving FinOps initiatives and cloud cost optimisation. Experience modernising legacy estates and delivering large-scale transformation programmes. Familiarity with event-driven architectures and platform automation. Experience working within highly regulated or secure environments.
16/07/2026
Full time
Senior Platform Engineer Salary: £100,000 - £110,000 base salary + benefits Build the platforms behind critical digital services Are you a Senior or Lead Platform Engineer who thrives on solving complex infrastructure challenges, building production-grade platforms, and shaping the engineering practices that enable teams to deliver at scale? We are looking for experienced Platform Engineers to help design, build and continuously improve critical digital services used across the Nation. You'll work on platforms that must remain secure, resilient and observable under significant demand, helping to modernise essential services that make a real difference. You'll join a highly collaborative engineering community where knowledge sharing, continuous improvement and technical excellence are at the heart of everything we do. The Opportunity As a Senior Platform Engineer, you'll play a key role in designing, building and evolving modern cloud platforms that enable engineering teams to deliver safely, rapidly and reliably. This is a hands-on technical leadership role where you'll combine deep engineering expertise with strategic influence. You'll work closely with delivery teams, architects and client stakeholders to define platform strategy, establish engineering standards and build reusable capabilities that improve the developer experience. You'll remain close to the technology, actively contributing to architecture, code, automation and operational improvements while mentoring and supporting other engineers. What You'll Be Doing You'll help design, build and operate the platforms that underpin mission-critical services, including: Designing secure, scalable multi-cloud landing zones, primarily within AWS. Building GitOps-driven platforms and Internal Developer Platforms (IDPs) that provide true self-service capabilities for engineering teams. Modernising legacy environments into cloud-native architectures. Creating unified observability solutions using OpenTelemetry, Prometheus, Grafana and modern APM tooling. Driving improvements in reliability, operability and platform performance through Site Reliability Engineering (SRE) practices. Implementing secure CI/CD pipelines with a strong focus on DevSecOps and software supply-chain integrity. Delivering event-driven automation and infrastructure capabilities. Driving FinOps initiatives and optimising cloud consumption across large-scale estates. Applying platform-as-a-product principles to create reusable, scalable engineering capabilities. Technical Leadership & Influence As a senior member of the engineering team, you will: Act as a senior technical authority, influencing platform strategy, architecture and governance. Lead technical workshops, architecture discussions and design reviews with stakeholders and delivery teams. Define and evolve engineering standards covering security, reliability, observability, CI/CD and operational excellence. Establish Service Level Objectives (SLOs), improve reliability and facilitate incident reviews and continuous improvement activities. Champion platform engineering best practices and help teams adopt them successfully. Coach and mentor engineers through pairing, technical leadership and knowledge sharing. Drive platform-as-a-product thinking, establishing golden paths and self-service capabilities that improve developer experience. Remain hands-on by contributing to designs, code reviews, automation and complex technical problem solving. What We're Looking For Essential Experience Significant experience designing, building and operating cloud-native platforms in enterprise or large-scale environments. Strong experience with AWS, including multi-account architectures and secure landing zones. Experience building and operating Kubernetes platforms, particularly EKS (AKS experience also beneficial). Expertise in Infrastructure as Code using tools such as Terraform. Strong understanding of CI/CD, DevSecOps and modern software delivery practices. Experience implementing observability solutions using tools such as OpenTelemetry, Prometheus and Grafana. Strong knowledge of cloud security, identity, networking and platform reliability. Experience leading technical discussions and influencing engineering direction across multiple teams. Proven ability to mentor engineers and provide technical leadership within multidisciplinary teams. Desirable Experience Experience with Azure and/or Google Cloud Platform. Experience building Internal Developer Platforms and self-service engineering capabilities. Knowledge of GitOps tooling such as Argo CD or Flux. Experience applying Site Reliability Engineering (SRE) practices. Experience driving FinOps initiatives and cloud cost optimisation. Experience modernising legacy estates and delivering large-scale transformation programmes. Familiarity with event-driven architectures and platform automation. Experience working within highly regulated or secure environments.
Valarian Technologies is a dual use technology company building critical tools to safeguard the future in an era of evolving global security challenges. We're rethinking security beyond traditional military domains, addressing asymmetric threats that impact our technological advantage, economic strength, and democratic institutions. We build Acra - the platform foundation for everything we do as a dual use technology company. The platform's name, rooted in the Greek word for citadel (or, fortress), reflects the design and purpose of our infrastructure agnostic secure enclaves: protecting critical data. Some of the government and commercial workflows include increased operational resiliency for mission critical systems and functions; enabling organizations to more quickly and widely adopt emerging technologies while ensuring the integrity of their intellectual property; information flow during disaster response scenarios, and zero trust / least privilege environments for M&A, attorney client privileged communications, etc. And we've only scratched the surface. At our core, we're driven by a shared mission and a belief in making a tangible impact on our world. Whether you join our London HQ or the wider global organisation, you'll be a part of collaborative, high performing teams, creating cutting edge software, platforms, and infrastructure. Role We're looking for a Senior Infrastructure Engineer to own and evolve the foundational infrastructure layer behind ACRA. This is a deep infrastructure role. You will work across Kubernetes, Linux, networking, storage, service to service communication, observability, security boundaries, and production operations. You will be responsible for the systems that everything else depends on: clusters, networks, storage layers, ingress and egress paths, runtime infrastructure, deployment foundations, and operational reliability. This role values depth, judgement, and operational correctness over short term churn. We are looking for someone with real production experience who can reason through complex infrastructure problems, understand failure modes, and make careful engineering decisions that improve the reliability and security of the whole platform. You should be comfortable operating close to the metal: debugging Kubernetes, understanding networking behaviour, reasoning about distributed storage, improving observability, and helping define the infrastructure patterns that ACRA will rely on as it scales. This is not a generic DevOps support role or internal IT role. You will be a critical engineer in the team responsible for the infrastructure foundations of a high trust platform. What you'll do: Design, build, operate, and improve Kubernetes based infrastructure for ACRA. Own core infrastructure plumbing across networking, storage, workload scheduling, service communication, ingress, egress, DNS, certificates, and cluster level security. Operate and debug production Kubernetes environments across GCP, on premise, bare metal, sovereign cloud, air gapped, and customer managed deployments. Work on multi cluster Kubernetes environments, cluster networking, network policy, service mesh, and secure service to service communication. Help operate and evolve distributed storage systems, including storage classes, CSI drivers, capacity planning, replication, recovery, and failure handling. Work with infrastructure technologies such as Kubernetes, Cilium, eBPF, Istio, Rook Ceph, Terraform, Argo CD, GitOps, Helm, Kustomize, and related CNCF tooling. Build infrastructure automation that improves repeatability, reliability, and operational safety. Improve observability across infrastructure layers, including metrics, logs, traces, alerting, dashboards, and operational runbooks. Investigate and resolve complex production issues across networking, storage, Kubernetes, Linux, and application infrastructure. Contribute to incident response, root cause analysis, capacity planning, disaster recovery, and production readiness. Define infrastructure standards, operational boundaries, security controls, and deployment patterns across the platform. Work closely with engineering teams to make services easier to deploy, operate, monitor, and debug. Contribute to the long term technical direction of the DevOps department and the infrastructure foundations of ACRA. What we are looking for: Strong experience in infrastructure engineering, platform engineering, DevOps, SRE, or production systems engineering. Deep hands on experience operating Kubernetes in production. Strong understanding of Linux systems, containers, networking, storage, and distributed infrastructure. Strong networking fundamentals, including TCP/IP, DNS, TLS, routing, load balancing, ingress, egress, network policy, and service discovery. Experience with Kubernetes networking, CNI, service mesh, or secure service to service communication. Experience debugging difficult infrastructure issues across clusters, nodes, pods, networks, storage, and workloads. Experience with infrastructure as code and GitOps workflows, especially Terraform, Argo CD, Helm, or Kustomize. Experience building reliable infrastructure for production environments where uptime, security, and operational clarity matter. Comfortable working across cloud, on premise, bare metal, restricted, or customer managed environments. Strong operational judgement, especially around reliability, resilience, failure domains, and production risk. Strong ownership mindset. You can take responsibility for critical systems and improve them over time. Clear communication skills. You can explain complex infrastructure problems to engineering and leadership without adding noise. Comfortable working in a startup environment with ambiguity, changing priorities, and broad technical responsibility. Nice to have: Experience with distributed storage systems, especially Rook Ceph or Ceph. Experience with Cilium, eBPF, cluster mesh, Istio, Envoy, or similar networking and service mesh technologies. Experience with multi cluster Kubernetes environments. Experience operating infrastructure in secure, regulated, sovereign, air gapped, defence, government, fintech, healthcare, critical infrastructure, or customer managed environments. Experience with zero trust architecture, workload isolation, identity aware networking, policy enforcement, or runtime security. Experience with secrets management tools such as Vault, OpenBao, SOPS, Sealed Secrets, External Secrets, or cloud native secret managers. Experience with supply chain security, including image signing, vulnerability scanning, SBOMs, artefact verification, admission control, or policy as code. Experience designing infrastructure where auditability, traceability, and operational evidence are first class requirements. Benefits: Equity - because you have the right to own what you're building. A competitive salary - because we value your unique skills. Employer pension contributions - because you deserve a secure future. Hybrid work setup - because everyone has different needs. Rewarding company retreats and meetups that respect your work/life balance - because we love getting to know each other! Life at Valarian Our culture is built on inclusivity, compassion and flexibility - we want everyone to be empowered to achieve their goals at Valarian. The world is changing, and so is our way of working. If you want to join us in the London office - it's quite nice! - you're welcome to; and if you'd prefer to work remotely, that's fine too. And as we build elegant solutions to simplify the complexities of business collaboration, we also simplify work life: contribute wherever and whenever allows you to be your best self. We trust each other to consistently raise the bar, and we challenge each other to continually reach new heights. Valarian Technologies Limited is an equal opportunity employer and welcomes applications from individuals regardless of race, colour, religion, sex, sexual orientation, gender, identity or expression, national origin, age, disability, genetic information, marital status, veteran, amnesty, or any other legally protected characteristic. We are committed to ensuring a fair and inclusive recruitment process and providing employment opportunities to all applicants. Decision recruitment, hiring, and employment are based solely on qualifications, skills, and experience relevant to the job requirements.
15/07/2026
Full time
Valarian Technologies is a dual use technology company building critical tools to safeguard the future in an era of evolving global security challenges. We're rethinking security beyond traditional military domains, addressing asymmetric threats that impact our technological advantage, economic strength, and democratic institutions. We build Acra - the platform foundation for everything we do as a dual use technology company. The platform's name, rooted in the Greek word for citadel (or, fortress), reflects the design and purpose of our infrastructure agnostic secure enclaves: protecting critical data. Some of the government and commercial workflows include increased operational resiliency for mission critical systems and functions; enabling organizations to more quickly and widely adopt emerging technologies while ensuring the integrity of their intellectual property; information flow during disaster response scenarios, and zero trust / least privilege environments for M&A, attorney client privileged communications, etc. And we've only scratched the surface. At our core, we're driven by a shared mission and a belief in making a tangible impact on our world. Whether you join our London HQ or the wider global organisation, you'll be a part of collaborative, high performing teams, creating cutting edge software, platforms, and infrastructure. Role We're looking for a Senior Infrastructure Engineer to own and evolve the foundational infrastructure layer behind ACRA. This is a deep infrastructure role. You will work across Kubernetes, Linux, networking, storage, service to service communication, observability, security boundaries, and production operations. You will be responsible for the systems that everything else depends on: clusters, networks, storage layers, ingress and egress paths, runtime infrastructure, deployment foundations, and operational reliability. This role values depth, judgement, and operational correctness over short term churn. We are looking for someone with real production experience who can reason through complex infrastructure problems, understand failure modes, and make careful engineering decisions that improve the reliability and security of the whole platform. You should be comfortable operating close to the metal: debugging Kubernetes, understanding networking behaviour, reasoning about distributed storage, improving observability, and helping define the infrastructure patterns that ACRA will rely on as it scales. This is not a generic DevOps support role or internal IT role. You will be a critical engineer in the team responsible for the infrastructure foundations of a high trust platform. What you'll do: Design, build, operate, and improve Kubernetes based infrastructure for ACRA. Own core infrastructure plumbing across networking, storage, workload scheduling, service communication, ingress, egress, DNS, certificates, and cluster level security. Operate and debug production Kubernetes environments across GCP, on premise, bare metal, sovereign cloud, air gapped, and customer managed deployments. Work on multi cluster Kubernetes environments, cluster networking, network policy, service mesh, and secure service to service communication. Help operate and evolve distributed storage systems, including storage classes, CSI drivers, capacity planning, replication, recovery, and failure handling. Work with infrastructure technologies such as Kubernetes, Cilium, eBPF, Istio, Rook Ceph, Terraform, Argo CD, GitOps, Helm, Kustomize, and related CNCF tooling. Build infrastructure automation that improves repeatability, reliability, and operational safety. Improve observability across infrastructure layers, including metrics, logs, traces, alerting, dashboards, and operational runbooks. Investigate and resolve complex production issues across networking, storage, Kubernetes, Linux, and application infrastructure. Contribute to incident response, root cause analysis, capacity planning, disaster recovery, and production readiness. Define infrastructure standards, operational boundaries, security controls, and deployment patterns across the platform. Work closely with engineering teams to make services easier to deploy, operate, monitor, and debug. Contribute to the long term technical direction of the DevOps department and the infrastructure foundations of ACRA. What we are looking for: Strong experience in infrastructure engineering, platform engineering, DevOps, SRE, or production systems engineering. Deep hands on experience operating Kubernetes in production. Strong understanding of Linux systems, containers, networking, storage, and distributed infrastructure. Strong networking fundamentals, including TCP/IP, DNS, TLS, routing, load balancing, ingress, egress, network policy, and service discovery. Experience with Kubernetes networking, CNI, service mesh, or secure service to service communication. Experience debugging difficult infrastructure issues across clusters, nodes, pods, networks, storage, and workloads. Experience with infrastructure as code and GitOps workflows, especially Terraform, Argo CD, Helm, or Kustomize. Experience building reliable infrastructure for production environments where uptime, security, and operational clarity matter. Comfortable working across cloud, on premise, bare metal, restricted, or customer managed environments. Strong operational judgement, especially around reliability, resilience, failure domains, and production risk. Strong ownership mindset. You can take responsibility for critical systems and improve them over time. Clear communication skills. You can explain complex infrastructure problems to engineering and leadership without adding noise. Comfortable working in a startup environment with ambiguity, changing priorities, and broad technical responsibility. Nice to have: Experience with distributed storage systems, especially Rook Ceph or Ceph. Experience with Cilium, eBPF, cluster mesh, Istio, Envoy, or similar networking and service mesh technologies. Experience with multi cluster Kubernetes environments. Experience operating infrastructure in secure, regulated, sovereign, air gapped, defence, government, fintech, healthcare, critical infrastructure, or customer managed environments. Experience with zero trust architecture, workload isolation, identity aware networking, policy enforcement, or runtime security. Experience with secrets management tools such as Vault, OpenBao, SOPS, Sealed Secrets, External Secrets, or cloud native secret managers. Experience with supply chain security, including image signing, vulnerability scanning, SBOMs, artefact verification, admission control, or policy as code. Experience designing infrastructure where auditability, traceability, and operational evidence are first class requirements. Benefits: Equity - because you have the right to own what you're building. A competitive salary - because we value your unique skills. Employer pension contributions - because you deserve a secure future. Hybrid work setup - because everyone has different needs. Rewarding company retreats and meetups that respect your work/life balance - because we love getting to know each other! Life at Valarian Our culture is built on inclusivity, compassion and flexibility - we want everyone to be empowered to achieve their goals at Valarian. The world is changing, and so is our way of working. If you want to join us in the London office - it's quite nice! - you're welcome to; and if you'd prefer to work remotely, that's fine too. And as we build elegant solutions to simplify the complexities of business collaboration, we also simplify work life: contribute wherever and whenever allows you to be your best self. We trust each other to consistently raise the bar, and we challenge each other to continually reach new heights. Valarian Technologies Limited is an equal opportunity employer and welcomes applications from individuals regardless of race, colour, religion, sex, sexual orientation, gender, identity or expression, national origin, age, disability, genetic information, marital status, veteran, amnesty, or any other legally protected characteristic. We are committed to ensuring a fair and inclusive recruitment process and providing employment opportunities to all applicants. Decision recruitment, hiring, and employment are based solely on qualifications, skills, and experience relevant to the job requirements.