Technical Product Manager - AI Cloud Infrastructure Era4 develops, owns and operates AI infrastructure across the UK, powered by renewable energy. Converting legacy industrial and energy sites into modern data centre facilities, Era4 is combining brownfield regeneration opportunities with cleaner, efficient, scalable compute capacity for healthcare, research, finance, enterprise, and public sector organisations. Role Summary We are seeking a Technical Product Manager - AI Cloud Infrastructure to join our fast scaling team. In this role, you will embed with engineering to act as the "First Customer," owning the continuous validation, reliability strategy, and technical documentation for our bare metal, VM, Kubernetes, and ML infrastructure. By treating testability as a core feature and shadowing real world workflows, you will ensure our compute platform handles the demands of advanced AI training and engineering workloads. This is an opportunity to join a mission led AI business that is redefining infrastructure, intelligence, and impact for enterprise customers. Key Responsibilities Execute integration testing in staging environments, work closely with the platform engineers to build repeatable test frameworks, and shadow internal and external AI infrastructure engineers to translate their real world usage patterns into automated in house test cases. Establish strict quality gates, performance SLOs, and scheduling benchmarks that our compute and orchestration services must pass before production deployment. Review, refine, and author technical guides, API documentation, and CLI guides, using them as the blueprint to test the platform exactly as an external engineer would. Partner with software and platform engineers to design robust validation suites, anticipating complex edge cases and structural failure modes across bare metal provisioning and Kubernetes cluster lifecycles. Technical familiarity with bare metal infrastructure (e.g., PXE booting, IPMI/Redfish), virtualization layers (e.g., KVM), and container orchestration (Kubernetes or similar). Track record designing comprehensive test strategies, validation frameworks, and acceptance criteria for highly technical cloud native, API, or infrastructure as a service (IaaS) products. Analyse infrastructure services, CLIs, and APIs from a developer's perspective to identify friction points, usability gaps, and reliability risks. Working knowledge of modern CI/CD pipelines, automated testing, and automation tooling (e.g., GitLab CI, GitHub Actions, Terraform, Ansible) to help engineering shape automated quality gates. Proven experience in a highly technical role embedded directly within a core infrastructure or platform engineering team. Desired Skills / Advantages Direct exposure to high performance computing (HPC) setups, large scale cluster scheduling (e.g., Slurm), or infrastructure optimized for heavy AI/ML training workloads. Experience using cloud observability, telemetry, and monitoring tools (e.g., Prometheus, Grafana, Datadog) to track and improve system reliability metrics. Experience writing or structuring technical documentation, API reference guides, and developer tutorials from scratch. Why Join Era4 You'll be joining a mission driven start up building critical national infrastructure, where operational excellence directly enables growth. This role offers high visibility with leadership, real autonomy, and the chance to shape how a next generation company operates at scale. Diversity & Inclusion Era4 is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. London, United Kingdom (Visit to office required)
22/07/2026
Full time
Technical Product Manager - AI Cloud Infrastructure Era4 develops, owns and operates AI infrastructure across the UK, powered by renewable energy. Converting legacy industrial and energy sites into modern data centre facilities, Era4 is combining brownfield regeneration opportunities with cleaner, efficient, scalable compute capacity for healthcare, research, finance, enterprise, and public sector organisations. Role Summary We are seeking a Technical Product Manager - AI Cloud Infrastructure to join our fast scaling team. In this role, you will embed with engineering to act as the "First Customer," owning the continuous validation, reliability strategy, and technical documentation for our bare metal, VM, Kubernetes, and ML infrastructure. By treating testability as a core feature and shadowing real world workflows, you will ensure our compute platform handles the demands of advanced AI training and engineering workloads. This is an opportunity to join a mission led AI business that is redefining infrastructure, intelligence, and impact for enterprise customers. Key Responsibilities Execute integration testing in staging environments, work closely with the platform engineers to build repeatable test frameworks, and shadow internal and external AI infrastructure engineers to translate their real world usage patterns into automated in house test cases. Establish strict quality gates, performance SLOs, and scheduling benchmarks that our compute and orchestration services must pass before production deployment. Review, refine, and author technical guides, API documentation, and CLI guides, using them as the blueprint to test the platform exactly as an external engineer would. Partner with software and platform engineers to design robust validation suites, anticipating complex edge cases and structural failure modes across bare metal provisioning and Kubernetes cluster lifecycles. Technical familiarity with bare metal infrastructure (e.g., PXE booting, IPMI/Redfish), virtualization layers (e.g., KVM), and container orchestration (Kubernetes or similar). Track record designing comprehensive test strategies, validation frameworks, and acceptance criteria for highly technical cloud native, API, or infrastructure as a service (IaaS) products. Analyse infrastructure services, CLIs, and APIs from a developer's perspective to identify friction points, usability gaps, and reliability risks. Working knowledge of modern CI/CD pipelines, automated testing, and automation tooling (e.g., GitLab CI, GitHub Actions, Terraform, Ansible) to help engineering shape automated quality gates. Proven experience in a highly technical role embedded directly within a core infrastructure or platform engineering team. Desired Skills / Advantages Direct exposure to high performance computing (HPC) setups, large scale cluster scheduling (e.g., Slurm), or infrastructure optimized for heavy AI/ML training workloads. Experience using cloud observability, telemetry, and monitoring tools (e.g., Prometheus, Grafana, Datadog) to track and improve system reliability metrics. Experience writing or structuring technical documentation, API reference guides, and developer tutorials from scratch. Why Join Era4 You'll be joining a mission driven start up building critical national infrastructure, where operational excellence directly enables growth. This role offers high visibility with leadership, real autonomy, and the chance to shape how a next generation company operates at scale. Diversity & Inclusion Era4 is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. London, United Kingdom (Visit to office required)
Carbon3ai Limited. is seeking a Technical Product Manager for AI Cloud Infrastructure in Greater London. You will own validation, reliability strategy, and documentation for complex AI workloads. The role demands technical familiarity with virtualization and container orchestration, alongside expertise in CI/CD pipelines and automated testing. Join our mission-driven team to build crucial national infrastructure, emphasizing diversity and inclusion.
22/07/2026
Full time
Carbon3ai Limited. is seeking a Technical Product Manager for AI Cloud Infrastructure in Greater London. You will own validation, reliability strategy, and documentation for complex AI workloads. The role demands technical familiarity with virtualization and container orchestration, alongside expertise in CI/CD pipelines and automated testing. Join our mission-driven team to build crucial national infrastructure, emphasizing diversity and inclusion.
Era4 develops, owns and operates AI infrastructure across the UK, powered by renewable energy. Converting legacy industrial and energy sites into modern data centre facilities, Era4 is combining brownfield regeneration opportunities with cleaner, efficient, scalable compute capacity for healthcare, research, finance, enterprise, and public sector organisations. We are seeking a Technical Writer to support the creation and maintenance of clear, structured documentation across our platform and infrastructure. You will work closely with engineering, platform, and product teams to translate complex technical concepts into accessible content for internal teams and customers. This role is hands on and delivery focused, ideal for someone who can bring structure to fast moving environments and improve how knowledge is captured and shared. Key Responsibilities: Create and maintain technical documentation including platform guides, onboarding materials, runbooks, and operational procedures. Work with engineering and platform teams to document systems, workflows, and APIs. Translate complex infrastructure and platform concepts into clear, user friendly content. Support customer facing documentation such as user guides and knowledge base articles. Maintain and improve documentation repositories (e.g., Confluence, Git based docs, Notion). Apply consistent standards, templates, and formatting across documentation. Keep documentation up to date as systems evolve, ensuring accuracy and usability. In depth experience as a Technical Writer in a software, cloud, or infrastructure environment. Strong ability to understand and explain technical systems (e.g., cloud platforms, Kubernetes, networking fundamentals). Experience working with engineers and product teams to produce documentation. Clear and concise writing style with strong attention to detail. Familiarity with documentation tools such as Markdown, Git, Confluence, or similar. Comfortable operating in a fast paced, evolving environment. One or more would be an advantage: Exposure to AI/ML platforms or GPU based infrastructure. Familiarity with Kubernetes or container based platforms. Experience documenting APIs or developer facing products. Understanding of data centre environments (compute, storage, networking). Experience in a startup or scaling organisation. Why Join Era4: You'll be joining a mission driven start up building critical national infrastructure, where operational excellence directly enables growth. This role offers high visibility with leadership, real autonomy, and the chance to shape how a next generation company operates at scale.
28/06/2026
Full time
Era4 develops, owns and operates AI infrastructure across the UK, powered by renewable energy. Converting legacy industrial and energy sites into modern data centre facilities, Era4 is combining brownfield regeneration opportunities with cleaner, efficient, scalable compute capacity for healthcare, research, finance, enterprise, and public sector organisations. We are seeking a Technical Writer to support the creation and maintenance of clear, structured documentation across our platform and infrastructure. You will work closely with engineering, platform, and product teams to translate complex technical concepts into accessible content for internal teams and customers. This role is hands on and delivery focused, ideal for someone who can bring structure to fast moving environments and improve how knowledge is captured and shared. Key Responsibilities: Create and maintain technical documentation including platform guides, onboarding materials, runbooks, and operational procedures. Work with engineering and platform teams to document systems, workflows, and APIs. Translate complex infrastructure and platform concepts into clear, user friendly content. Support customer facing documentation such as user guides and knowledge base articles. Maintain and improve documentation repositories (e.g., Confluence, Git based docs, Notion). Apply consistent standards, templates, and formatting across documentation. Keep documentation up to date as systems evolve, ensuring accuracy and usability. In depth experience as a Technical Writer in a software, cloud, or infrastructure environment. Strong ability to understand and explain technical systems (e.g., cloud platforms, Kubernetes, networking fundamentals). Experience working with engineers and product teams to produce documentation. Clear and concise writing style with strong attention to detail. Familiarity with documentation tools such as Markdown, Git, Confluence, or similar. Comfortable operating in a fast paced, evolving environment. One or more would be an advantage: Exposure to AI/ML platforms or GPU based infrastructure. Familiarity with Kubernetes or container based platforms. Experience documenting APIs or developer facing products. Understanding of data centre environments (compute, storage, networking). Experience in a startup or scaling organisation. Why Join Era4: You'll be joining a mission driven start up building critical national infrastructure, where operational excellence directly enables growth. This role offers high visibility with leadership, real autonomy, and the chance to shape how a next generation company operates at scale.
A leading tech company is seeking a Technical Writer to create and maintain structured documentation for its platform and infrastructure. The ideal candidate will translate complex technical concepts into accessible content and work closely with engineering and product teams. Responsibilities include creating user guides and operational procedures, improving documentation formats, and ensuring accuracy. This role offers the chance to shape documentation practices in a dynamic environment, perfect for those passionate about technology and effective communication.
27/06/2026
Full time
A leading tech company is seeking a Technical Writer to create and maintain structured documentation for its platform and infrastructure. The ideal candidate will translate complex technical concepts into accessible content and work closely with engineering and product teams. Responsibilities include creating user guides and operational procedures, improving documentation formats, and ensuring accuracy. This role offers the chance to shape documentation practices in a dynamic environment, perfect for those passionate about technology and effective communication.
Carbon3ai Limited. in the United Kingdom is seeking Automation Engineers to build a modern AI Platform Operations function from scratch. You will develop workflows, build tooling for performance enhancement, and maintain observability platforms within a hybrid work environment. Ideal candidates will have strong Python skills, experience with automation and observability tools like Prometheus and Grafana, and a passion for operational excellence. Diversity and inclusion are emphasized in our work culture.
27/06/2026
Full time
Carbon3ai Limited. in the United Kingdom is seeking Automation Engineers to build a modern AI Platform Operations function from scratch. You will develop workflows, build tooling for performance enhancement, and maintain observability platforms within a hybrid work environment. Ideal candidates will have strong Python skills, experience with automation and observability tools like Prometheus and Grafana, and a passion for operational excellence. Diversity and inclusion are emphasized in our work culture.
Overview Era4 develops, owns and operates AI infrastructure across the UK, powered by renewable energy. Converting legacy industrial and energy sites into modern data centre facilities, Era4 is combining brownfield regeneration opportunities with cleaner, efficient, scalable compute capacity for healthcare, research, finance, enterprise, and public sector organisations. Role Summary We are seeking Automation Engineers who sit at the intersection of Site Reliability Engineering and modern AI driven operations. Embedded within Era4's engineering led Operations Centre, this role exists to build a modern AI Platform Operations function from scratch, designing tooling, and agentic workflows. No legacy to deal with. Key Responsibilities Runbook Automation & Agent Development: Build agentic, executable workflows capable of triaging, diagnosing, and where appropriate autonomously remediating known failure patterns. Build and maintain LLM backed agents targeting the observability stack, ITSM platform, and infrastructure APIs (e.g. DCIM, IPAM, hypervisor layers). Develop auditable client focused automations, for client interactions and workflows, with appropriate controls. Develop safe, auditable automation with appropriate controls for higher risk platform actions. Operational Tooling & Self Service Enablement: Build internal tooling that empowers engineers and service desk analysts: CLI utilities, ChatOps integrations (Slack/Teams bots), status dashboards, and self service automation hooks. Reduce dependency on DevSecOps and engineering teams for routine operational tasks through automation. Maintain and contribute a library of automation assets, agent prompts, and runbook as code artefacts, version controlled and peer reviewed. Develop the automation layer around monitoring and event management: alert suppression logic, enrichment pipelines, correlation rules, and alert to ticket integrations. Continuously tune signal to noise ratios across monitoring tooling (Prometheus, Mimir, Grafana, or equivalent) to improve situational awareness. Design and implement event correlation and deduplication logic to reduce alert storms and improve incident context. Identify common operational patterns and tasks as candidates for automation; maintain and prioritise a toil reduction backlog. Participate in post incident reviews and translate findings into updated automation, runbooks, or agent logic. Contribute to the evolution of Era4's operational standards, tooling architecture, and agent framework. Operational: Prior experience in an SRE, Senior Operations, or Platform Engineering environment, with exposure to on call operations and incident management processes. Experience in converting narrative runbooks into executable automation or codified decision trees. Understanding of ITIL aligned incident and change management principles and ITSM tooling. Technical - Core Element Strong Python development skills, including scripting for automation, API integration, and data processing. Hands on experience with observability and monitoring platforms: Prometheus, Grafana, Mimir, or equivalent. Experience integrating with ITSM platforms (ServiceNow, Halo, Jira Service Management, or similar) via API. Solid understanding of event driven architectures, message queues, and webhook based automation patterns. Strong understanding of managing GPU infrastructure in production, key signals and metrics and the automation of workflows. Familiarity with Infrastructure as Code principles and cloud native environments (Kubernetes, Terraform, or similar). Comfort operating in an API first environment, integrating agents with infrastructure APIs, DCIM, IPAM, and hypervisor control planes. One or more would be an advantage Exposure to data centre or colocation operations, particularly high density compute or GPU infrastructure environments. Experience with ChatOps tooling: building Slack or Microsoft Teams bots for operational workflows. Familiarity with DCIM platforms and telemetry pipelines (power, thermal, network). Knowledge of OpenTelemetry, distributed tracing, or log aggregation platforms (Loki, ELK, Splunk). Contributions to open source observability or automation tooling. Experience in a start up or scale up environment where tooling is built from scratch. Why Join Era4 You'll be joining a mission driven start up building critical national infrastructure, where operational excellence directly enables growth. This role offers high visibility with leadership, real autonomy, and the chance to shape how a next generation company operates at scale. Diversity & Inclusion Era4 is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. United Kingdom - Hybrid (Visit to office / site required)
27/06/2026
Full time
Overview Era4 develops, owns and operates AI infrastructure across the UK, powered by renewable energy. Converting legacy industrial and energy sites into modern data centre facilities, Era4 is combining brownfield regeneration opportunities with cleaner, efficient, scalable compute capacity for healthcare, research, finance, enterprise, and public sector organisations. Role Summary We are seeking Automation Engineers who sit at the intersection of Site Reliability Engineering and modern AI driven operations. Embedded within Era4's engineering led Operations Centre, this role exists to build a modern AI Platform Operations function from scratch, designing tooling, and agentic workflows. No legacy to deal with. Key Responsibilities Runbook Automation & Agent Development: Build agentic, executable workflows capable of triaging, diagnosing, and where appropriate autonomously remediating known failure patterns. Build and maintain LLM backed agents targeting the observability stack, ITSM platform, and infrastructure APIs (e.g. DCIM, IPAM, hypervisor layers). Develop auditable client focused automations, for client interactions and workflows, with appropriate controls. Develop safe, auditable automation with appropriate controls for higher risk platform actions. Operational Tooling & Self Service Enablement: Build internal tooling that empowers engineers and service desk analysts: CLI utilities, ChatOps integrations (Slack/Teams bots), status dashboards, and self service automation hooks. Reduce dependency on DevSecOps and engineering teams for routine operational tasks through automation. Maintain and contribute a library of automation assets, agent prompts, and runbook as code artefacts, version controlled and peer reviewed. Develop the automation layer around monitoring and event management: alert suppression logic, enrichment pipelines, correlation rules, and alert to ticket integrations. Continuously tune signal to noise ratios across monitoring tooling (Prometheus, Mimir, Grafana, or equivalent) to improve situational awareness. Design and implement event correlation and deduplication logic to reduce alert storms and improve incident context. Identify common operational patterns and tasks as candidates for automation; maintain and prioritise a toil reduction backlog. Participate in post incident reviews and translate findings into updated automation, runbooks, or agent logic. Contribute to the evolution of Era4's operational standards, tooling architecture, and agent framework. Operational: Prior experience in an SRE, Senior Operations, or Platform Engineering environment, with exposure to on call operations and incident management processes. Experience in converting narrative runbooks into executable automation or codified decision trees. Understanding of ITIL aligned incident and change management principles and ITSM tooling. Technical - Core Element Strong Python development skills, including scripting for automation, API integration, and data processing. Hands on experience with observability and monitoring platforms: Prometheus, Grafana, Mimir, or equivalent. Experience integrating with ITSM platforms (ServiceNow, Halo, Jira Service Management, or similar) via API. Solid understanding of event driven architectures, message queues, and webhook based automation patterns. Strong understanding of managing GPU infrastructure in production, key signals and metrics and the automation of workflows. Familiarity with Infrastructure as Code principles and cloud native environments (Kubernetes, Terraform, or similar). Comfort operating in an API first environment, integrating agents with infrastructure APIs, DCIM, IPAM, and hypervisor control planes. One or more would be an advantage Exposure to data centre or colocation operations, particularly high density compute or GPU infrastructure environments. Experience with ChatOps tooling: building Slack or Microsoft Teams bots for operational workflows. Familiarity with DCIM platforms and telemetry pipelines (power, thermal, network). Knowledge of OpenTelemetry, distributed tracing, or log aggregation platforms (Loki, ELK, Splunk). Contributions to open source observability or automation tooling. Experience in a start up or scale up environment where tooling is built from scratch. Why Join Era4 You'll be joining a mission driven start up building critical national infrastructure, where operational excellence directly enables growth. This role offers high visibility with leadership, real autonomy, and the chance to shape how a next generation company operates at scale. Diversity & Inclusion Era4 is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. United Kingdom - Hybrid (Visit to office / site required)