it job board logo
  • Home
  • Find IT Jobs
  • Register CV
  • Career Advice
  • Contact us
  • Employers
    • Register as Employer
    • Pricing Plans
  • Recruiting? Post a job
  • Sign in
  • Sign up
  • Home
  • Find IT Jobs
  • Register CV
  • Career Advice
  • Contact us
  • Employers
    • Register as Employer
    • Pricing Plans
Sorry, that job is no longer available. Here are some results that may be similar to the job you were looking for.

73 jobs found

Email me jobs like this
Refine Search
Current Search
lead site reliability engineer sre aws azure
Inspire People
Lead Site Reliability Engineer (SRE Squad Lead)
Inspire People Cardiff, South Glamorgan
Become part of a mission-driven digital team helping to build reliable, secure and scalable digital services that support economic growth across the UK. The Department for Business and Trade (DBT), in partnership with Inspire People, is seeking a Senior SRE Squad Lead with experience leading and developing engineers, strong DevOps and Site Reliability Engineering expertise, cloud platform experience, infrastructure-as-code capability and a passion for building resilient distributed systems. Based in London, Cardiff, Darlington, Edinburgh, Belfast, Birmingham or Salford, this permanent Grade 7 opportunity offers hybrid working, flexible working patterns and a salary of £63,824 to £80,158 (London £67,547 to £83,778,) depending on location and technical skills assessed at interview. Shape Reliable Digital Services at the Department for Business and Trade The Department for Business and Trade has a clear mission: to grow the economy by helping businesses invest, grow and export, creating jobs and opportunities across the UK. DBT's Digital, Data and Technology directorate develops and operates the tools and services that enable this mission. As a Senior SRE Squad Lead, you will play a key role in leading engineers while remaining hands-on in the design, delivery and continuous improvement of reliable, secure and scalable platform services that underpin critical digital products and services. As a Senior SRE Squad Lead you will: Lead and support a team of Site Reliability Engineers, setting clear direction while fostering an inclusive, collaborative and high-performing team culture. Build strong working relationships with product, delivery and architecture colleagues to ensure platform services meet business and user needs. Provide technical leadership across DevOps and SRE practices, guiding teams to adopt approaches that support reliability, sustainability and continuous improvement. Coach, mentor and support engineers across DDaT, contributing to a supportive and diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets in a pragmatic and user-focused way. Support the development and continual improvement of CI/CD pipelines to enable safe, frequent and low-risk delivery of changes. Oversee live service reliability, supporting teams through incident and problem management while encouraging a learning-focused, blameless culture. Ensure security, resilience and compliance considerations are understood and embedded into engineering practices. Essential skills for the Senior SRE Squad Lead include: Experience leading, supporting and developing engineers, including line management or strong mentoring experience. Strong communication skills, with the ability to explain technical concepts clearly and build effective relationships with a range of stakeholders. Experience of working with cloud platforms such as AWS, Azure or Google Cloud, applying modern DevOps and SRE practices. Experience designing and delivering infrastructure-as-code solutions using tools such as Terraform, CloudFormation or similar. Ability to write clean, maintainable and well-tested code in at least one programming language. Experience designing, operating and improving distributed systems, with a focus on reliability, performance and user impact. In return, you can expect a flexible working culture, including: Flexible hybrid working, typically 2-3 days per week in the office. Full-time, part-time and flexible working options. A Civil Service pension with an average employer contribution of 27%. Annual leave starting at 25 days, rising to 30 days with service. Three paid volunteering days every year. Learning and development tailored to your role. Access to professional qualifications, certifications and technical training. An inclusive culture that encourages learning, collaboration and continuous improvement. Employee benefits including cycle-to-work and wider Civil Service benefits Why Join DBT? This is an opportunity to combine technical leadership with hands-on engineering in a modern cloud environment. You'll work alongside experienced SRE and DevOps professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore new ideas, while making a tangible impact on services used across the Department for Business and Trade. This role requires SC clearance. DBT's requirement for SC clearance is to have been present in the UK for at least 3 of the last 5 years. Failure to meet this requirement will result in your application being rejected and your offer withdrawn. If you're an experienced Site Reliability Engineering leader who wants to remain close to technology while helping engineers thrive and delivering services that matter, apply today via Inspire People. JBRP1_UKTJ
16/07/2026
Full time
Become part of a mission-driven digital team helping to build reliable, secure and scalable digital services that support economic growth across the UK. The Department for Business and Trade (DBT), in partnership with Inspire People, is seeking a Senior SRE Squad Lead with experience leading and developing engineers, strong DevOps and Site Reliability Engineering expertise, cloud platform experience, infrastructure-as-code capability and a passion for building resilient distributed systems. Based in London, Cardiff, Darlington, Edinburgh, Belfast, Birmingham or Salford, this permanent Grade 7 opportunity offers hybrid working, flexible working patterns and a salary of £63,824 to £80,158 (London £67,547 to £83,778,) depending on location and technical skills assessed at interview. Shape Reliable Digital Services at the Department for Business and Trade The Department for Business and Trade has a clear mission: to grow the economy by helping businesses invest, grow and export, creating jobs and opportunities across the UK. DBT's Digital, Data and Technology directorate develops and operates the tools and services that enable this mission. As a Senior SRE Squad Lead, you will play a key role in leading engineers while remaining hands-on in the design, delivery and continuous improvement of reliable, secure and scalable platform services that underpin critical digital products and services. As a Senior SRE Squad Lead you will: Lead and support a team of Site Reliability Engineers, setting clear direction while fostering an inclusive, collaborative and high-performing team culture. Build strong working relationships with product, delivery and architecture colleagues to ensure platform services meet business and user needs. Provide technical leadership across DevOps and SRE practices, guiding teams to adopt approaches that support reliability, sustainability and continuous improvement. Coach, mentor and support engineers across DDaT, contributing to a supportive and diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets in a pragmatic and user-focused way. Support the development and continual improvement of CI/CD pipelines to enable safe, frequent and low-risk delivery of changes. Oversee live service reliability, supporting teams through incident and problem management while encouraging a learning-focused, blameless culture. Ensure security, resilience and compliance considerations are understood and embedded into engineering practices. Essential skills for the Senior SRE Squad Lead include: Experience leading, supporting and developing engineers, including line management or strong mentoring experience. Strong communication skills, with the ability to explain technical concepts clearly and build effective relationships with a range of stakeholders. Experience of working with cloud platforms such as AWS, Azure or Google Cloud, applying modern DevOps and SRE practices. Experience designing and delivering infrastructure-as-code solutions using tools such as Terraform, CloudFormation or similar. Ability to write clean, maintainable and well-tested code in at least one programming language. Experience designing, operating and improving distributed systems, with a focus on reliability, performance and user impact. In return, you can expect a flexible working culture, including: Flexible hybrid working, typically 2-3 days per week in the office. Full-time, part-time and flexible working options. A Civil Service pension with an average employer contribution of 27%. Annual leave starting at 25 days, rising to 30 days with service. Three paid volunteering days every year. Learning and development tailored to your role. Access to professional qualifications, certifications and technical training. An inclusive culture that encourages learning, collaboration and continuous improvement. Employee benefits including cycle-to-work and wider Civil Service benefits Why Join DBT? This is an opportunity to combine technical leadership with hands-on engineering in a modern cloud environment. You'll work alongside experienced SRE and DevOps professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore new ideas, while making a tangible impact on services used across the Department for Business and Trade. This role requires SC clearance. DBT's requirement for SC clearance is to have been present in the UK for at least 3 of the last 5 years. Failure to meet this requirement will result in your application being rejected and your offer withdrawn. If you're an experienced Site Reliability Engineering leader who wants to remain close to technology while helping engineers thrive and delivering services that matter, apply today via Inspire People. JBRP1_UKTJ
Senior Site Reliability Engineer
Trades Workforce Solutions Wokingham, Berkshire
Senior Site Reliability Engineer (SRE) Location: Wokingham (2 days/week onsite) Type: Inside IR35 Rate: Up to £70.00 per hour (DOE) We're looking for a Senior Site Reliability Engineer (SRE) to lead efforts in maintaining the reliability, performance, and scalability of mission-critical platforms and services. This role is ideal for someone who thrives at the intersection of software engineering, infrastructure, automation, and incident response. You'll be instrumental in defining and implementing the standards and systems that keep applications running smoothly across cloud and hybrid environments-including OpenShift clusters. What You'll Be Responsible For As a Senior SRE, you will: Ensure high availability, performance, and latency of critical systems across Azure, AWS, and OpenShift. Design and implement robust observability systems (logging, monitoring, alerting) to detect and resolve issues proactively. Lead and evolve incident management processes-runbooks, comms, postmortems, and root cause analysis. Define and monitor SLIs, SLOs, and error budgets to balance innovation with stability. Automate manual processes through infrastructure-as-code, scripting, and modern CI/CD pipelines. Mentor engineering teams on best practices for deployment, reliability, scalability, and incident preparedness. Support and scale OpenShift-based containerized applications, including upgrade strategies, patching, and workload optimization. Core Responsibilities Operations & Incident Management Act as the senior escalation point for outages and critical incidents. Lead post-incident reviews and implement long-term remediation plans. Communicate platform health and risk posture to stakeholders at all levels. Engineering & Automation Build and improve CI/CD pipelines using tools like Azure DevOps, GitHub Actions, Jenkins, and GitLab. Design scalable, fault-tolerant infrastructure with IaC tools (Terraform, Bicep). Create internal tools and automation to accelerate development and reduce operational toil. Strategic & Advisory Architect cloud and container infrastructure, with a focus on OpenShift, Kubernetes, and hybrid deployments. Collaborate with engineering, architecture, and security teams to embed reliability into the SDLC. Promote advanced deployment strategies (blue-green, canary, rolling updates) and rollback readiness. Drive a culture of reliability, observability, and operational excellence across engineering teams. Technical Environment Hands-on experience with many of the following is expected: Cloud & Containers: Azure, AWS, OpenShift, Kubernetes, Docker, App Services, IaaS (EC2, VMs) CI/CD & Automation: Terraform, Bicep, Azure DevOps, Jenkins, GitHub Actions, GitLab Observability: Prometheus, Grafana, Datadog, ELK, Splunk, Application Insights, CloudWatch Languages & Scripting: Python, C#, Bash, PowerShell Networking: DNS, SSL/TLS, load balancing, WAF, proxies, CDN, Azure App Gateway Databases: MSSQL, PostgreSQL, MongoDB, CosmosDB, DynamoDB OS & Systems: Windows, Linux, Nginx, IIS Ideal Candidate Profile 5+ years of experience in SRE, DevOps, or production engineering roles. Expertise operating in high-availability, fast-paced production environments. Solid engineering foundation with experience reading and writing production code. Hands-on experience deploying, supporting, and scaling OpenShift environments. Proven track record of leading incident responses and improving system reliability. Strong collaboration and mentoring abilities across infrastructure, development, and security teams. What You'll Bring Ability to balance operational risk with engineering velocity. Strong communication skills across technical and non-technical audiences. A passion for automating everything and eliminating manual work. A mindset of ownership, continuous improvement, and technical leadership. Ready to make reliability your legacy? If you're a senior SRE with OpenShift experience and a drive to solve complex operational challenges, we'd love to hear from you.
16/07/2026
Full time
Senior Site Reliability Engineer (SRE) Location: Wokingham (2 days/week onsite) Type: Inside IR35 Rate: Up to £70.00 per hour (DOE) We're looking for a Senior Site Reliability Engineer (SRE) to lead efforts in maintaining the reliability, performance, and scalability of mission-critical platforms and services. This role is ideal for someone who thrives at the intersection of software engineering, infrastructure, automation, and incident response. You'll be instrumental in defining and implementing the standards and systems that keep applications running smoothly across cloud and hybrid environments-including OpenShift clusters. What You'll Be Responsible For As a Senior SRE, you will: Ensure high availability, performance, and latency of critical systems across Azure, AWS, and OpenShift. Design and implement robust observability systems (logging, monitoring, alerting) to detect and resolve issues proactively. Lead and evolve incident management processes-runbooks, comms, postmortems, and root cause analysis. Define and monitor SLIs, SLOs, and error budgets to balance innovation with stability. Automate manual processes through infrastructure-as-code, scripting, and modern CI/CD pipelines. Mentor engineering teams on best practices for deployment, reliability, scalability, and incident preparedness. Support and scale OpenShift-based containerized applications, including upgrade strategies, patching, and workload optimization. Core Responsibilities Operations & Incident Management Act as the senior escalation point for outages and critical incidents. Lead post-incident reviews and implement long-term remediation plans. Communicate platform health and risk posture to stakeholders at all levels. Engineering & Automation Build and improve CI/CD pipelines using tools like Azure DevOps, GitHub Actions, Jenkins, and GitLab. Design scalable, fault-tolerant infrastructure with IaC tools (Terraform, Bicep). Create internal tools and automation to accelerate development and reduce operational toil. Strategic & Advisory Architect cloud and container infrastructure, with a focus on OpenShift, Kubernetes, and hybrid deployments. Collaborate with engineering, architecture, and security teams to embed reliability into the SDLC. Promote advanced deployment strategies (blue-green, canary, rolling updates) and rollback readiness. Drive a culture of reliability, observability, and operational excellence across engineering teams. Technical Environment Hands-on experience with many of the following is expected: Cloud & Containers: Azure, AWS, OpenShift, Kubernetes, Docker, App Services, IaaS (EC2, VMs) CI/CD & Automation: Terraform, Bicep, Azure DevOps, Jenkins, GitHub Actions, GitLab Observability: Prometheus, Grafana, Datadog, ELK, Splunk, Application Insights, CloudWatch Languages & Scripting: Python, C#, Bash, PowerShell Networking: DNS, SSL/TLS, load balancing, WAF, proxies, CDN, Azure App Gateway Databases: MSSQL, PostgreSQL, MongoDB, CosmosDB, DynamoDB OS & Systems: Windows, Linux, Nginx, IIS Ideal Candidate Profile 5+ years of experience in SRE, DevOps, or production engineering roles. Expertise operating in high-availability, fast-paced production environments. Solid engineering foundation with experience reading and writing production code. Hands-on experience deploying, supporting, and scaling OpenShift environments. Proven track record of leading incident responses and improving system reliability. Strong collaboration and mentoring abilities across infrastructure, development, and security teams. What You'll Bring Ability to balance operational risk with engineering velocity. Strong communication skills across technical and non-technical audiences. A passion for automating everything and eliminating manual work. A mindset of ownership, continuous improvement, and technical leadership. Ready to make reliability your legacy? If you're a senior SRE with OpenShift experience and a drive to solve complex operational challenges, we'd love to hear from you.
SF Partners
Senior Platform Engineer
SF Partners
Senior Platform Engineer Salary: £100,000 - £110,000 base salary + benefits Build the platforms behind critical digital services Are you a Senior or Lead Platform Engineer who thrives on solving complex infrastructure challenges, building production-grade platforms, and shaping the engineering practices that enable teams to deliver at scale? We are looking for experienced Platform Engineers to help design, build and continuously improve critical digital services used across the Nation. You'll work on platforms that must remain secure, resilient and observable under significant demand, helping to modernise essential services that make a real difference. You'll join a highly collaborative engineering community where knowledge sharing, continuous improvement and technical excellence are at the heart of everything we do. The Opportunity As a Senior Platform Engineer, you'll play a key role in designing, building and evolving modern cloud platforms that enable engineering teams to deliver safely, rapidly and reliably. This is a hands-on technical leadership role where you'll combine deep engineering expertise with strategic influence. You'll work closely with delivery teams, architects and client stakeholders to define platform strategy, establish engineering standards and build reusable capabilities that improve the developer experience. You'll remain close to the technology, actively contributing to architecture, code, automation and operational improvements while mentoring and supporting other engineers. What You'll Be Doing You'll help design, build and operate the platforms that underpin mission-critical services, including: Designing secure, scalable multi-cloud landing zones, primarily within AWS. Building GitOps-driven platforms and Internal Developer Platforms (IDPs) that provide true self-service capabilities for engineering teams. Modernising legacy environments into cloud-native architectures. Creating unified observability solutions using OpenTelemetry, Prometheus, Grafana and modern APM tooling. Driving improvements in reliability, operability and platform performance through Site Reliability Engineering (SRE) practices. Implementing secure CI/CD pipelines with a strong focus on DevSecOps and software supply-chain integrity. Delivering event-driven automation and infrastructure capabilities. Driving FinOps initiatives and optimising cloud consumption across large-scale estates. Applying platform-as-a-product principles to create reusable, scalable engineering capabilities. Technical Leadership & Influence As a senior member of the engineering team, you will: Act as a senior technical authority, influencing platform strategy, architecture and governance. Lead technical workshops, architecture discussions and design reviews with stakeholders and delivery teams. Define and evolve engineering standards covering security, reliability, observability, CI/CD and operational excellence. Establish Service Level Objectives (SLOs), improve reliability and facilitate incident reviews and continuous improvement activities. Champion platform engineering best practices and help teams adopt them successfully. Coach and mentor engineers through pairing, technical leadership and knowledge sharing. Drive platform-as-a-product thinking, establishing golden paths and self-service capabilities that improve developer experience. Remain hands-on by contributing to designs, code reviews, automation and complex technical problem solving. What We're Looking For Essential Experience Significant experience designing, building and operating cloud-native platforms in enterprise or large-scale environments. Strong experience with AWS, including multi-account architectures and secure landing zones. Experience building and operating Kubernetes platforms, particularly EKS (AKS experience also beneficial). Expertise in Infrastructure as Code using tools such as Terraform. Strong understanding of CI/CD, DevSecOps and modern software delivery practices. Experience implementing observability solutions using tools such as OpenTelemetry, Prometheus and Grafana. Strong knowledge of cloud security, identity, networking and platform reliability. Experience leading technical discussions and influencing engineering direction across multiple teams. Proven ability to mentor engineers and provide technical leadership within multidisciplinary teams. Desirable Experience Experience with Azure and/or Google Cloud Platform. Experience building Internal Developer Platforms and self-service engineering capabilities. Knowledge of GitOps tooling such as Argo CD or Flux. Experience applying Site Reliability Engineering (SRE) practices. Experience driving FinOps initiatives and cloud cost optimisation. Experience modernising legacy estates and delivering large-scale transformation programmes. Familiarity with event-driven architectures and platform automation. Experience working within highly regulated or secure environments.
16/07/2026
Full time
Senior Platform Engineer Salary: £100,000 - £110,000 base salary + benefits Build the platforms behind critical digital services Are you a Senior or Lead Platform Engineer who thrives on solving complex infrastructure challenges, building production-grade platforms, and shaping the engineering practices that enable teams to deliver at scale? We are looking for experienced Platform Engineers to help design, build and continuously improve critical digital services used across the Nation. You'll work on platforms that must remain secure, resilient and observable under significant demand, helping to modernise essential services that make a real difference. You'll join a highly collaborative engineering community where knowledge sharing, continuous improvement and technical excellence are at the heart of everything we do. The Opportunity As a Senior Platform Engineer, you'll play a key role in designing, building and evolving modern cloud platforms that enable engineering teams to deliver safely, rapidly and reliably. This is a hands-on technical leadership role where you'll combine deep engineering expertise with strategic influence. You'll work closely with delivery teams, architects and client stakeholders to define platform strategy, establish engineering standards and build reusable capabilities that improve the developer experience. You'll remain close to the technology, actively contributing to architecture, code, automation and operational improvements while mentoring and supporting other engineers. What You'll Be Doing You'll help design, build and operate the platforms that underpin mission-critical services, including: Designing secure, scalable multi-cloud landing zones, primarily within AWS. Building GitOps-driven platforms and Internal Developer Platforms (IDPs) that provide true self-service capabilities for engineering teams. Modernising legacy environments into cloud-native architectures. Creating unified observability solutions using OpenTelemetry, Prometheus, Grafana and modern APM tooling. Driving improvements in reliability, operability and platform performance through Site Reliability Engineering (SRE) practices. Implementing secure CI/CD pipelines with a strong focus on DevSecOps and software supply-chain integrity. Delivering event-driven automation and infrastructure capabilities. Driving FinOps initiatives and optimising cloud consumption across large-scale estates. Applying platform-as-a-product principles to create reusable, scalable engineering capabilities. Technical Leadership & Influence As a senior member of the engineering team, you will: Act as a senior technical authority, influencing platform strategy, architecture and governance. Lead technical workshops, architecture discussions and design reviews with stakeholders and delivery teams. Define and evolve engineering standards covering security, reliability, observability, CI/CD and operational excellence. Establish Service Level Objectives (SLOs), improve reliability and facilitate incident reviews and continuous improvement activities. Champion platform engineering best practices and help teams adopt them successfully. Coach and mentor engineers through pairing, technical leadership and knowledge sharing. Drive platform-as-a-product thinking, establishing golden paths and self-service capabilities that improve developer experience. Remain hands-on by contributing to designs, code reviews, automation and complex technical problem solving. What We're Looking For Essential Experience Significant experience designing, building and operating cloud-native platforms in enterprise or large-scale environments. Strong experience with AWS, including multi-account architectures and secure landing zones. Experience building and operating Kubernetes platforms, particularly EKS (AKS experience also beneficial). Expertise in Infrastructure as Code using tools such as Terraform. Strong understanding of CI/CD, DevSecOps and modern software delivery practices. Experience implementing observability solutions using tools such as OpenTelemetry, Prometheus and Grafana. Strong knowledge of cloud security, identity, networking and platform reliability. Experience leading technical discussions and influencing engineering direction across multiple teams. Proven ability to mentor engineers and provide technical leadership within multidisciplinary teams. Desirable Experience Experience with Azure and/or Google Cloud Platform. Experience building Internal Developer Platforms and self-service engineering capabilities. Knowledge of GitOps tooling such as Argo CD or Flux. Experience applying Site Reliability Engineering (SRE) practices. Experience driving FinOps initiatives and cloud cost optimisation. Experience modernising legacy estates and delivering large-scale transformation programmes. Familiarity with event-driven architectures and platform automation. Experience working within highly regulated or secure environments.
Infrastructure Engineer
Rex Technologies GmbH
About Marex Marex Group plc (NASDAQ: MRX) is a diversified global financial services platform providing essential liquidity, market access and infrastructure services to clients across energy, commodities and financial markets. The group provides comprehensive breadth and depth of coverage across four core services: clearing, agency and execution, market making, and hedging and investment solutions. It has a leading franchise in many major metals, energy and agricultural products, with access to 60 exchanges. The group provides access to the world's major commodity markets, covering a broad range of clients that include some of the largest commodity producers, consumers and traders, banks, hedge funds and asset managers. With more than 40 offices worldwide, the group has over 3,000 employees across Europe, Asia and the Americas. For more information visit Vacancy: VN2959 Department Description Marex has unique access across markets with significant share globally both on and off exchange. The depth of knowledge amongst its teams and divisions provides its customers with clear advantage, and its technology led service provides access to all major exchanges, order flow management via screen, voice and DMA, plus award winning data, insights, and analytics. The Technology Department delivers differentiation, scalability, and security for the business. Reporting to the COO, Technology provides digital tools, software services and infrastructure globally to all business groups. Software development and support teams work in agile 'streams' aligned to specific business areas. Our other teams work enterprise wide to provide critical services including our global service desk, network and system infrastructure, IT operations, security, enterprise architecture and design. IT runs our enterprise wide services to end users and actively manages the firm's infrastructure and data to provide and accelerate business value. A Data team enables the firm to leverage data to increase productivity and improve business decisions, as well as maintain data compliance. Our Infrastructure group delivers operations and engineering across Infrastructure Operations, Network, Communications, Endpoint and Platform Engineering. Infrastructure Operations is responsible for the stability, scalability, and reliability of Marex's global technology platforms. This team ensures seamless day to day operations across cloud and on premises environments, including Azure, AWS, VMware, Citrix, storage services, and Office 365 products. By leveraging automation, proactive monitoring, and best practices in site reliability engineering (SRE), Infrastructure Operations minimises downtime, optimises performance, and enhances security. The team collaborates closely with business units to drive innovation, streamline processes, and support Marex's strategic technology initiatives. Role Summary The Infrastructure Engineer is a hands on automation engineer who applies strong coding and programming skills to build, improve, and upgrade Marex's infrastructure. Writing production grade PowerShell and Python, and authoring Infrastructure as Code with Terraform and Ansible, the successful candidate automates the configuration, deployment, and management of the technologies Marex runs every day - Office 365, VMware, and services across AWS and Azure. This is a delivery focused engineering role: most of the work is writing code, building CI/CD pipelines, and replacing manual effort with reliable, repeatable automation. The Infrastructure Engineer collaborates with the global IT team across APAC, Europe, and North America to engineer, monitor, and maintain infrastructure systems, ensuring maximum uptime and optimal service delivery to the business. As a senior member of the team, the successful candidate sets the technical standard for coding and automation practices and mentors others, while remaining hands on as the primary focus of the role. We are seeking an engineer who treats infrastructure as software and has a genuine passion for automation. The ideal candidate is fluent in PowerShell and Python and uses Terraform, Ansible, and BitBucket Pipelines to automate configuration, deployment, monitoring, and troubleshooting across our Office 365, VMware, AWS, and Azure estate, engineering away manual, repetitive work wherever it exists. Responsibilities Role specific Design, write, and maintain automation code in PowerShell and Python to configure, deploy, and manage infrastructure across Office 365, VMware, AWS, and Azure. Author and maintain Infrastructure as Code using Terraform and Ansible, applying software engineering best practices such as version control, code review, and modular, reusable design. Build and maintain CI/CD pipelines in BitBucket Pipelines to test and deploy infrastructure changes automatically and safely. Apply SRE principles by codifying reliability, self healing, and observability directly into automated solutions. Set coding and automation standards for the team and mentor colleagues while remaining hands on; identify manual, repetitive tasks and replace them with reliable, well documented automation. Engineer and automate platform migration and modernisation work across Azure, AWS, VMware, Citrix, storage services, and Office 365. Instrument systems and build automated monitoring and observability using tools such as Splunk, Grafana, and Opsgenie. Participate in on call rotations and incident response, automating detection and remediation to minimise downtime. Embed security and governance controls into automation code and pipelines so that changes are compliant by default. Be available on weekends to carry out infrastructure changes. Participate in Agile working methodology, delivering projects in sprints. Follow the change control process for making changes to the production environment. All staff Ensuring compliance with the company's regulatory requirements under the FCA, NFA, AMF, AFM, MAS. Adhere to the operational risk framework for your role ensuring that all regulatory or company determined parameters are complied with. Role model for demonstrating highest level standards of integrity and conduct and reflecting Company Values. At all times complying with the FCA's Code of Conduct To ensure that you are fully aware of and adhere to internal policies that relate to you, your role or any other activities for which you have any level of responsibility To report any breaches of policy to Compliance and/ or your supervisor as required To escape risk events immediately To provide input to risk management processes, as required. Competencies, Skills, Experience & Qualifications Skills and Experience Essential: 8+ years of hands on experience in infrastructure automation and engineering roles, with a strong track record of solving infrastructure problems through code. Advanced scripting and programming ability in PowerShell and Python, applied to real infrastructure automation and tooling. Expertise in Infrastructure as Code and CI/CD pipelines using Terraform, Ansible, and BitBucket Pipelines. Deep knowledge of the platforms these skills are applied to: Office 365, VMware, and services across AWS and Azure, plus Citrix and storage services. Strong understanding of observability, logging, and monitoring frameworks. Experience working in Agile environments and driving automation at scale. Ability to set technical standards and mentor others on coding and automation best practice, while remaining hands on. candidates outside of this range will also be considered Desirable A degree in Computer Science, Engineering, or a related field. Relevant certifications (AWS, Azure, Office 365, VMware) are a plus. Competencies A collaborative team player, approachable, self sufficient and influences a positive work environment. Demonstrates curiosity. Resilient in a challenging, fast paced environment Ability to take a high level of responsibility in a fast pace and high volume environment. Excels at building relationships, networking and influencing others. Strategic collaborator with insight and agility, able to anticipate future challenges, ensuring operational effectiveness Conduct Rules Act with integrity Act with due skill, care and diligence Be open and cooperative with the FCA, the PRA and other regulators Pay due regard to the interests of customers and treat them fairly Observe proper standard of market conduct Act to deliver good outcomes for retail customers Company Values Be collaborative - by working together across the organisation, we foster teamwork, can better respond to challenges and successfully deliver for our clients Act with integrity - we pride ourselves on our honesty and high ethical standards. We apply these values when working with all our clients, colleagues and other stakeholders Be adaptable and entrepreneurial - we embrace change as markets evolve to constantly increase our efficiency and create innovative solutions for our clients. We are interested in the world around us and inquisitive about understanding the challenges and opportunities our clients face. Be respectful - how we treat each other, and our clients says everything about who we are. We always act respectfully and treat people fairly in everything we do. Nurture talent . click apply for full job details
16/07/2026
Full time
About Marex Marex Group plc (NASDAQ: MRX) is a diversified global financial services platform providing essential liquidity, market access and infrastructure services to clients across energy, commodities and financial markets. The group provides comprehensive breadth and depth of coverage across four core services: clearing, agency and execution, market making, and hedging and investment solutions. It has a leading franchise in many major metals, energy and agricultural products, with access to 60 exchanges. The group provides access to the world's major commodity markets, covering a broad range of clients that include some of the largest commodity producers, consumers and traders, banks, hedge funds and asset managers. With more than 40 offices worldwide, the group has over 3,000 employees across Europe, Asia and the Americas. For more information visit Vacancy: VN2959 Department Description Marex has unique access across markets with significant share globally both on and off exchange. The depth of knowledge amongst its teams and divisions provides its customers with clear advantage, and its technology led service provides access to all major exchanges, order flow management via screen, voice and DMA, plus award winning data, insights, and analytics. The Technology Department delivers differentiation, scalability, and security for the business. Reporting to the COO, Technology provides digital tools, software services and infrastructure globally to all business groups. Software development and support teams work in agile 'streams' aligned to specific business areas. Our other teams work enterprise wide to provide critical services including our global service desk, network and system infrastructure, IT operations, security, enterprise architecture and design. IT runs our enterprise wide services to end users and actively manages the firm's infrastructure and data to provide and accelerate business value. A Data team enables the firm to leverage data to increase productivity and improve business decisions, as well as maintain data compliance. Our Infrastructure group delivers operations and engineering across Infrastructure Operations, Network, Communications, Endpoint and Platform Engineering. Infrastructure Operations is responsible for the stability, scalability, and reliability of Marex's global technology platforms. This team ensures seamless day to day operations across cloud and on premises environments, including Azure, AWS, VMware, Citrix, storage services, and Office 365 products. By leveraging automation, proactive monitoring, and best practices in site reliability engineering (SRE), Infrastructure Operations minimises downtime, optimises performance, and enhances security. The team collaborates closely with business units to drive innovation, streamline processes, and support Marex's strategic technology initiatives. Role Summary The Infrastructure Engineer is a hands on automation engineer who applies strong coding and programming skills to build, improve, and upgrade Marex's infrastructure. Writing production grade PowerShell and Python, and authoring Infrastructure as Code with Terraform and Ansible, the successful candidate automates the configuration, deployment, and management of the technologies Marex runs every day - Office 365, VMware, and services across AWS and Azure. This is a delivery focused engineering role: most of the work is writing code, building CI/CD pipelines, and replacing manual effort with reliable, repeatable automation. The Infrastructure Engineer collaborates with the global IT team across APAC, Europe, and North America to engineer, monitor, and maintain infrastructure systems, ensuring maximum uptime and optimal service delivery to the business. As a senior member of the team, the successful candidate sets the technical standard for coding and automation practices and mentors others, while remaining hands on as the primary focus of the role. We are seeking an engineer who treats infrastructure as software and has a genuine passion for automation. The ideal candidate is fluent in PowerShell and Python and uses Terraform, Ansible, and BitBucket Pipelines to automate configuration, deployment, monitoring, and troubleshooting across our Office 365, VMware, AWS, and Azure estate, engineering away manual, repetitive work wherever it exists. Responsibilities Role specific Design, write, and maintain automation code in PowerShell and Python to configure, deploy, and manage infrastructure across Office 365, VMware, AWS, and Azure. Author and maintain Infrastructure as Code using Terraform and Ansible, applying software engineering best practices such as version control, code review, and modular, reusable design. Build and maintain CI/CD pipelines in BitBucket Pipelines to test and deploy infrastructure changes automatically and safely. Apply SRE principles by codifying reliability, self healing, and observability directly into automated solutions. Set coding and automation standards for the team and mentor colleagues while remaining hands on; identify manual, repetitive tasks and replace them with reliable, well documented automation. Engineer and automate platform migration and modernisation work across Azure, AWS, VMware, Citrix, storage services, and Office 365. Instrument systems and build automated monitoring and observability using tools such as Splunk, Grafana, and Opsgenie. Participate in on call rotations and incident response, automating detection and remediation to minimise downtime. Embed security and governance controls into automation code and pipelines so that changes are compliant by default. Be available on weekends to carry out infrastructure changes. Participate in Agile working methodology, delivering projects in sprints. Follow the change control process for making changes to the production environment. All staff Ensuring compliance with the company's regulatory requirements under the FCA, NFA, AMF, AFM, MAS. Adhere to the operational risk framework for your role ensuring that all regulatory or company determined parameters are complied with. Role model for demonstrating highest level standards of integrity and conduct and reflecting Company Values. At all times complying with the FCA's Code of Conduct To ensure that you are fully aware of and adhere to internal policies that relate to you, your role or any other activities for which you have any level of responsibility To report any breaches of policy to Compliance and/ or your supervisor as required To escape risk events immediately To provide input to risk management processes, as required. Competencies, Skills, Experience & Qualifications Skills and Experience Essential: 8+ years of hands on experience in infrastructure automation and engineering roles, with a strong track record of solving infrastructure problems through code. Advanced scripting and programming ability in PowerShell and Python, applied to real infrastructure automation and tooling. Expertise in Infrastructure as Code and CI/CD pipelines using Terraform, Ansible, and BitBucket Pipelines. Deep knowledge of the platforms these skills are applied to: Office 365, VMware, and services across AWS and Azure, plus Citrix and storage services. Strong understanding of observability, logging, and monitoring frameworks. Experience working in Agile environments and driving automation at scale. Ability to set technical standards and mentor others on coding and automation best practice, while remaining hands on. candidates outside of this range will also be considered Desirable A degree in Computer Science, Engineering, or a related field. Relevant certifications (AWS, Azure, Office 365, VMware) are a plus. Competencies A collaborative team player, approachable, self sufficient and influences a positive work environment. Demonstrates curiosity. Resilient in a challenging, fast paced environment Ability to take a high level of responsibility in a fast pace and high volume environment. Excels at building relationships, networking and influencing others. Strategic collaborator with insight and agility, able to anticipate future challenges, ensuring operational effectiveness Conduct Rules Act with integrity Act with due skill, care and diligence Be open and cooperative with the FCA, the PRA and other regulators Pay due regard to the interests of customers and treat them fairly Observe proper standard of market conduct Act to deliver good outcomes for retail customers Company Values Be collaborative - by working together across the organisation, we foster teamwork, can better respond to challenges and successfully deliver for our clients Act with integrity - we pride ourselves on our honesty and high ethical standards. We apply these values when working with all our clients, colleagues and other stakeholders Be adaptable and entrepreneurial - we embrace change as markets evolve to constantly increase our efficiency and create innovative solutions for our clients. We are interested in the world around us and inquisitive about understanding the challenges and opportunities our clients face. Be respectful - how we treat each other, and our clients says everything about who we are. We always act respectfully and treat people fairly in everything we do. Nurture talent . click apply for full job details
Remote Senior Site Reliability Engineer Manager (Remote)
Remotestar
Job description RemoteStar is looking to hire a Senior Site Reliability Engineering Manager on behalf of our client based in the UK with a fully remote work policy. About Client The client building, the B2B marketplace for diamonds. It's an industry-leading B2B diamond and gemstones marketplace, connecting jewelry retailers to gemstone supplies They have a presence in London, Hong Kong, Amsterdam, and as well in Mumbai and now in New York in 2001. About the role As the SRE Manager, you will play a critical role in ensuring the reliability, scalability, and performance of our infrastructure and services through both direct technical contribution along with team building and management. Take full ownership of the production estate from both a technical and process perspective. Provide a consistent smooth operation of live systems and drive all on-call support issues. Design and operate a new incident tracking process to ensure root causes are found and remediated in a timely fashion by the development team. Create and maintain high end monitoring and automation tooling. Drive automation initiatives to streamline operational workflows and improve efficiency. Develop and maintain tools, scripts, and dashboards to monitor system health, performance, and reliability. Build a first class SRE team. Through a combination of leading by example, coaching and mentoring, mould the team would want to have around you. Provide leadership and guidance to the SRE team, fostering a culture of collaboration, innovation, and continuous improvement. Responsibilities Proven experience in a senior or lead SRE role, with a strong track record of building and maintaining highly reliable infrastructure and services. Expertise in incident management, including incident response, resolution, and post-mortem analysis. Proficiency in monitoring, alerting, and observability tools such as Prometheus, Grafana, ELK stack or Datadog. Experience with cloud platforms such as AWS, Azure, or GCP, including infrastructure as code tools like Terraform or CloudFormation. Strong scripting and automation skills, with proficiency in languages such as Python, Bash, or Go. Excellent communication and collaboration skills, with the ability to work effectively with cross-functional teams in a remote environment. Demonstrated leadership capabilities, with a passion for mentoring and developing team members. What they offer Dynamic working environment in an extremely fast-growing company Work in an international environment Work in a pleasant environment with very little hierarchy Intellectually challenging, play a massive role in client's success and scalability Flexible working hours
16/07/2026
Full time
Job description RemoteStar is looking to hire a Senior Site Reliability Engineering Manager on behalf of our client based in the UK with a fully remote work policy. About Client The client building, the B2B marketplace for diamonds. It's an industry-leading B2B diamond and gemstones marketplace, connecting jewelry retailers to gemstone supplies They have a presence in London, Hong Kong, Amsterdam, and as well in Mumbai and now in New York in 2001. About the role As the SRE Manager, you will play a critical role in ensuring the reliability, scalability, and performance of our infrastructure and services through both direct technical contribution along with team building and management. Take full ownership of the production estate from both a technical and process perspective. Provide a consistent smooth operation of live systems and drive all on-call support issues. Design and operate a new incident tracking process to ensure root causes are found and remediated in a timely fashion by the development team. Create and maintain high end monitoring and automation tooling. Drive automation initiatives to streamline operational workflows and improve efficiency. Develop and maintain tools, scripts, and dashboards to monitor system health, performance, and reliability. Build a first class SRE team. Through a combination of leading by example, coaching and mentoring, mould the team would want to have around you. Provide leadership and guidance to the SRE team, fostering a culture of collaboration, innovation, and continuous improvement. Responsibilities Proven experience in a senior or lead SRE role, with a strong track record of building and maintaining highly reliable infrastructure and services. Expertise in incident management, including incident response, resolution, and post-mortem analysis. Proficiency in monitoring, alerting, and observability tools such as Prometheus, Grafana, ELK stack or Datadog. Experience with cloud platforms such as AWS, Azure, or GCP, including infrastructure as code tools like Terraform or CloudFormation. Strong scripting and automation skills, with proficiency in languages such as Python, Bash, or Go. Excellent communication and collaboration skills, with the ability to work effectively with cross-functional teams in a remote environment. Demonstrated leadership capabilities, with a passion for mentoring and developing team members. What they offer Dynamic working environment in an extremely fast-growing company Work in an international environment Work in a pleasant environment with very little hierarchy Intellectually challenging, play a massive role in client's success and scalability Flexible working hours
Staff Implementation Engineer
Mosaic.tech
Harness is the AI Software Delivery Platform company, led by technologist and entrepreneur Jyoti Bansal (founder of AppDynamics, acquired by Cisco for $3.7B). Harness has raised approximately $570M in funding and is valued at $5.5B, backed by leading investors including Goldman Sachs, Menlo Ventures, IVP, Unusual Ventures, Citi Ventures, and more. As AI accelerates code creation, the real bottleneck has shifted to everything after the code - testing, deployments, application security, reliability, compliance, and cost optimization. Harness brings AI and automation to this "outer loop," helping teams ship software faster while maintaining security and governance throughout the entire software delivery lifecycle. Powered by Harness AI and the Software Delivery Knowledge Graph, the Harness Platform applies deep context and intelligent automation across the software delivery lifecycle with governance and policy-driven controls embedded throughout the platform. Over the past year, Harness powered over 185M deployments, 82M builds, 18T flag evaluations, 8M security scans, 9.1B optimized tests, 3T protected API calls, and helped manage $2.8B in cloud spend - enabling customers like United Airlines, Morningstar, and Choice Hotels to accelerate releases by up to 75%, reduce cloud costs by up to 60%, and achieve 10x DevOps efficiency. With a global team across 26 offices and 27 countries, Harness is shaping the future of AI software delivery - and we're looking for exceptional talent to help us move even faster. Position Summary In this role, you will be working with internal and external stakeholders to architect, design and implement FinOps, IaC, and DB solutions for enterprise customers. You will have an opportunity to work with Harness Engineering and various customer functions, such as DevOps, SRE, Cloud, Finance and Engineering Analytics teams. You will develop best practices and automations to streamline Harness platform deployments in the most efficient, scalable, repeatable and reliable manner possible. We're a high-growth company on a once-in-a-lifetime journey to revolutionize engineering deployment tools & continuous delivery. In this role, you will partner closely with enterprise customers across EMEA to architect, design, and implement DevSecOps, FinOps, and Engineering Excellence solutions using the Harness platform. You will collaborate with Harness Engineering, Sales Engineering, and Customer Success, as well as customer teams including DevOps, SRE, Platform Engineering, Cloud Infrastructure, Finance, and Engineering Leadership. Your mission is to help customers modernize and scale their CI/CD practices in a way that is secure, cost efficient, reliable, and aligned with business outcomes. This is a hands on, customer facing role in a high growth global SaaS company, supporting organizations as they evolve their cloud native and Kubernetes adoption across AWS, GCP, and Azure. About the Role Engage directly with customer technical teams to assess and understand existing processes and DevSecOps, CI/CD pipelines and governance models. Architect and implement an optimized Harness setup for integration, scale, and repeatability. Interface with the Customer's Executive and Leadership teams to understand the technical goals and business objectives related to their processes, design their Harness implementation to best fit those requirements, and correlate the technical success criteria to the business requirements. Provide positive anecdotes from each engagement, craft best practices around Customer implementations, convert them into automation and create reference patterns. Document and implement processes and solutions that are employed for onboarding success for the purpose of internal enablement. Contribute to the product design, assist in the Harness Community, and for building out of an advanced technical knowledge base. Consult on DevSecOps, cloud native CI/CD and Kubernetes. Interact with customers on a professional, meaningful and technically deep level. Define current state vs. target state architecture and a phased onboarding and adoption roadmap spanning CI/CD, IaC, security and reliability. Provide forward looking technical direction aligned with both customer needs and platform evolution. Author and standardize design patterns (reusable pipeline templates, deployment strategies, golden paths) and integration blueprints to SCM, cloud, artifact registries, secret managers, identity, ITSM, and observability. Guide and oversee execution across implementation engineers and delivery partners. Drive onboarding and adoption across multiple teams and business units. Deliver hands on implementation across CI/CD, infrastructure as code, supply chain security, the internal developer portal, database DevOps, and site reliability. Partner with Customer Success and account teams to drive adoption, growth, and long term customer success. Work closely with Pre sales and Post sales teams to ensure that Harness customers are successful and experience a high level of customer satisfaction with the Harness solution. About You BA/BS degree in CS or Computer Engineering related field with 8+ years in consulting roles / software delivery / DevSecOps / Platform Engineering. You have a strong consulting attitude and a commercial acumen in delivering custom solutions to customers of all sizes, stakeholder management, workshop facilitation, and customer facing communication are part of your DNA. Your written and verbal communication are exceptional - You are able to present to a platform engineer team as well as a CxO with equal confidence. You have experience driving engineering culture change in large enterprises - such as onboarding and adoption of new tooling, processes, and ways of working at scale. You are a results driven individual with a hunger for accomplishing in fast paced environments and a knack for optimizing processes. You have a proven ability to work with cross functional, distributed teams. You are a perpetual learner, thrive in a team setting, enjoy sharing your experience and solutions, consistently pursuing excellence and success in all your tasks, detail oriented and analytical, with excellent written and verbal communication skills. You are familiar with architecture frameworks (TOGAF or equivalent) and Well Architected reviews (AWS/Azure/GCP). You have proven hands on experience leading migrations from incumbent tooling to a modern, consolidated delivery platform - including continuous delivery (e.g., IBM UrbanCode Deploy / UCD), continuous integration (e.g., Jenkins), and infrastructure as code management (e.g., Terraform Cloud / HCP Terraform). Willingness to travel up to 25%. Last but not least, you have deep hands on experience in at least 3 of the areas listed below, with working knowledge of the rest: CI/CD & pipelines: GitLab CI/CD, GitHub Actions, Jenkins, Azure DevOps, CircleCI, GitOps principles with Argo CD/Workflows, Tekton, Buildkite, Bamboo, TeamCity, AWS CodePipeline. Policy as Code & governance: OPA/Rego, HashiCorp Sentinel, Kyverno. SSO & identity: Okta, Azure AD/Entra ID, OneLogin, Keycloak, LDAP, SAML/OIDC. Infrastructure as Code (IaC): Terraform, OpenTofu, Pulumi, AWS CloudFormation, Azure ARM/Bicep, Crossplane, Ansible - plus IaC orchestration/management such as Terraform Cloud / HCP Terraform, Spacelift, env0, Scalr, Terragrunt, Atlantis, AWS Proton. IaC security scanning: Checkov, Wiz, Snyk IaC, Checkmarx, Terrascan. Application security testing: SonarQube, Checkmarx, Veracode, Semgrep, Trivy, Grype, GitGuardian. Internal developer portals & platform engineering: Backstage, Cortex, Roadie, Kratix, Port, OpsLevel, Atlassian Compass, Configure8, Humanitec - with hands on catalog modeling, golden path scaffolding, and scorecards. Environment management & deployment: Argo CD, Flux, Spinnaker, Octopus Deploy, IBM UrbanCode Deploy, Codefresh; ephemer / on demand environments (Bunnyshell, Qovery, Garden, Okteto, Signadot); configuration management with Ansible, Puppet, Chef, Saltstack in conjunction with Helm and Kustomize usage. Database DevOps: Liquibase, Flyway (incl. Redgate Flyway), Bytebase, Atlas (Ariga), Sqitch, SchemaHero, Redgate SQL Change Automation, DBmaestro. Site reliability & incident response (incl. AI assisted): incident management and postmortems (PagerDuty, Opsgenie, incident.io, Rootly, FireHydrant, Blameless); AIOps / AI SRE (Resolve.ai, Cleric, Traversal, Neubird, Datadog Bits AI, BigPanda, Moogsoft). Observability: Datadog, Prometheus/Grafana, New Relic, Splunk, Dynatrace, Honeycomb, OpenTelemetry. Containers, cloud & foundations: Docker and Kubernetes (Helm, manifests, operators) on EKS/AKS/GKE; depth in at least one of AWS/Azure/GCP (IAM, networking, managed services); secrets management (HashiCorp Vault, AWS Secrets Manager, Azure Key Vault, GCP Secret Manager, CyberArk). Scripting, languages & OS: Linux, Bash, Python (or Go), PowerShell. ITSM: ServiceNow, Jira. Work Location UK, London. Travel is roughly 25% throughout the calendar year. What You Will Have at Harness Comprehensive healthcare benefits Flexible work schedule Paid Time Off and Parental Leave Monthly, quarterly, and annual social and team building events . click apply for full job details
15/07/2026
Full time
Harness is the AI Software Delivery Platform company, led by technologist and entrepreneur Jyoti Bansal (founder of AppDynamics, acquired by Cisco for $3.7B). Harness has raised approximately $570M in funding and is valued at $5.5B, backed by leading investors including Goldman Sachs, Menlo Ventures, IVP, Unusual Ventures, Citi Ventures, and more. As AI accelerates code creation, the real bottleneck has shifted to everything after the code - testing, deployments, application security, reliability, compliance, and cost optimization. Harness brings AI and automation to this "outer loop," helping teams ship software faster while maintaining security and governance throughout the entire software delivery lifecycle. Powered by Harness AI and the Software Delivery Knowledge Graph, the Harness Platform applies deep context and intelligent automation across the software delivery lifecycle with governance and policy-driven controls embedded throughout the platform. Over the past year, Harness powered over 185M deployments, 82M builds, 18T flag evaluations, 8M security scans, 9.1B optimized tests, 3T protected API calls, and helped manage $2.8B in cloud spend - enabling customers like United Airlines, Morningstar, and Choice Hotels to accelerate releases by up to 75%, reduce cloud costs by up to 60%, and achieve 10x DevOps efficiency. With a global team across 26 offices and 27 countries, Harness is shaping the future of AI software delivery - and we're looking for exceptional talent to help us move even faster. Position Summary In this role, you will be working with internal and external stakeholders to architect, design and implement FinOps, IaC, and DB solutions for enterprise customers. You will have an opportunity to work with Harness Engineering and various customer functions, such as DevOps, SRE, Cloud, Finance and Engineering Analytics teams. You will develop best practices and automations to streamline Harness platform deployments in the most efficient, scalable, repeatable and reliable manner possible. We're a high-growth company on a once-in-a-lifetime journey to revolutionize engineering deployment tools & continuous delivery. In this role, you will partner closely with enterprise customers across EMEA to architect, design, and implement DevSecOps, FinOps, and Engineering Excellence solutions using the Harness platform. You will collaborate with Harness Engineering, Sales Engineering, and Customer Success, as well as customer teams including DevOps, SRE, Platform Engineering, Cloud Infrastructure, Finance, and Engineering Leadership. Your mission is to help customers modernize and scale their CI/CD practices in a way that is secure, cost efficient, reliable, and aligned with business outcomes. This is a hands on, customer facing role in a high growth global SaaS company, supporting organizations as they evolve their cloud native and Kubernetes adoption across AWS, GCP, and Azure. About the Role Engage directly with customer technical teams to assess and understand existing processes and DevSecOps, CI/CD pipelines and governance models. Architect and implement an optimized Harness setup for integration, scale, and repeatability. Interface with the Customer's Executive and Leadership teams to understand the technical goals and business objectives related to their processes, design their Harness implementation to best fit those requirements, and correlate the technical success criteria to the business requirements. Provide positive anecdotes from each engagement, craft best practices around Customer implementations, convert them into automation and create reference patterns. Document and implement processes and solutions that are employed for onboarding success for the purpose of internal enablement. Contribute to the product design, assist in the Harness Community, and for building out of an advanced technical knowledge base. Consult on DevSecOps, cloud native CI/CD and Kubernetes. Interact with customers on a professional, meaningful and technically deep level. Define current state vs. target state architecture and a phased onboarding and adoption roadmap spanning CI/CD, IaC, security and reliability. Provide forward looking technical direction aligned with both customer needs and platform evolution. Author and standardize design patterns (reusable pipeline templates, deployment strategies, golden paths) and integration blueprints to SCM, cloud, artifact registries, secret managers, identity, ITSM, and observability. Guide and oversee execution across implementation engineers and delivery partners. Drive onboarding and adoption across multiple teams and business units. Deliver hands on implementation across CI/CD, infrastructure as code, supply chain security, the internal developer portal, database DevOps, and site reliability. Partner with Customer Success and account teams to drive adoption, growth, and long term customer success. Work closely with Pre sales and Post sales teams to ensure that Harness customers are successful and experience a high level of customer satisfaction with the Harness solution. About You BA/BS degree in CS or Computer Engineering related field with 8+ years in consulting roles / software delivery / DevSecOps / Platform Engineering. You have a strong consulting attitude and a commercial acumen in delivering custom solutions to customers of all sizes, stakeholder management, workshop facilitation, and customer facing communication are part of your DNA. Your written and verbal communication are exceptional - You are able to present to a platform engineer team as well as a CxO with equal confidence. You have experience driving engineering culture change in large enterprises - such as onboarding and adoption of new tooling, processes, and ways of working at scale. You are a results driven individual with a hunger for accomplishing in fast paced environments and a knack for optimizing processes. You have a proven ability to work with cross functional, distributed teams. You are a perpetual learner, thrive in a team setting, enjoy sharing your experience and solutions, consistently pursuing excellence and success in all your tasks, detail oriented and analytical, with excellent written and verbal communication skills. You are familiar with architecture frameworks (TOGAF or equivalent) and Well Architected reviews (AWS/Azure/GCP). You have proven hands on experience leading migrations from incumbent tooling to a modern, consolidated delivery platform - including continuous delivery (e.g., IBM UrbanCode Deploy / UCD), continuous integration (e.g., Jenkins), and infrastructure as code management (e.g., Terraform Cloud / HCP Terraform). Willingness to travel up to 25%. Last but not least, you have deep hands on experience in at least 3 of the areas listed below, with working knowledge of the rest: CI/CD & pipelines: GitLab CI/CD, GitHub Actions, Jenkins, Azure DevOps, CircleCI, GitOps principles with Argo CD/Workflows, Tekton, Buildkite, Bamboo, TeamCity, AWS CodePipeline. Policy as Code & governance: OPA/Rego, HashiCorp Sentinel, Kyverno. SSO & identity: Okta, Azure AD/Entra ID, OneLogin, Keycloak, LDAP, SAML/OIDC. Infrastructure as Code (IaC): Terraform, OpenTofu, Pulumi, AWS CloudFormation, Azure ARM/Bicep, Crossplane, Ansible - plus IaC orchestration/management such as Terraform Cloud / HCP Terraform, Spacelift, env0, Scalr, Terragrunt, Atlantis, AWS Proton. IaC security scanning: Checkov, Wiz, Snyk IaC, Checkmarx, Terrascan. Application security testing: SonarQube, Checkmarx, Veracode, Semgrep, Trivy, Grype, GitGuardian. Internal developer portals & platform engineering: Backstage, Cortex, Roadie, Kratix, Port, OpsLevel, Atlassian Compass, Configure8, Humanitec - with hands on catalog modeling, golden path scaffolding, and scorecards. Environment management & deployment: Argo CD, Flux, Spinnaker, Octopus Deploy, IBM UrbanCode Deploy, Codefresh; ephemer / on demand environments (Bunnyshell, Qovery, Garden, Okteto, Signadot); configuration management with Ansible, Puppet, Chef, Saltstack in conjunction with Helm and Kustomize usage. Database DevOps: Liquibase, Flyway (incl. Redgate Flyway), Bytebase, Atlas (Ariga), Sqitch, SchemaHero, Redgate SQL Change Automation, DBmaestro. Site reliability & incident response (incl. AI assisted): incident management and postmortems (PagerDuty, Opsgenie, incident.io, Rootly, FireHydrant, Blameless); AIOps / AI SRE (Resolve.ai, Cleric, Traversal, Neubird, Datadog Bits AI, BigPanda, Moogsoft). Observability: Datadog, Prometheus/Grafana, New Relic, Splunk, Dynatrace, Honeycomb, OpenTelemetry. Containers, cloud & foundations: Docker and Kubernetes (Helm, manifests, operators) on EKS/AKS/GKE; depth in at least one of AWS/Azure/GCP (IAM, networking, managed services); secrets management (HashiCorp Vault, AWS Secrets Manager, Azure Key Vault, GCP Secret Manager, CyberArk). Scripting, languages & OS: Linux, Bash, Python (or Go), PowerShell. ITSM: ServiceNow, Jira. Work Location UK, London. Travel is roughly 25% throughout the calendar year. What You Will Have at Harness Comprehensive healthcare benefits Flexible work schedule Paid Time Off and Parental Leave Monthly, quarterly, and annual social and team building events . click apply for full job details
Hays Technology
Platform Engineer (GCP)
Hays Technology City, Manchester
Prestigious opportunity for a talented and experienced Platform Engineer to join a rapidly growing digital engineering team delivering cutting edge solutions across a diverse portfolio of clients.This is hands on applying your deep engineering and architectural expertise to design, build, and evolve scalable platform solutions that drive digital transformation within some of the UK's most exciting organisations.Collaborating closely with highly skilled, cross functional teams, you will work alongside engineers, architects, and product specialists to push boundaries and deliver impactful, high quality outcomes. Key responsibilities: Design, build, and maintain robust cloud platforms and CI/CD pipelines Contribute across the full software development lifecycle Provide technical leadership and input into system and architecture design Collaborate with cross functional teams to deliver scalable, reliable solutions. Write clean, well documented code and contribute to technical documentation Proactively monitor, troubleshoot, and resolve production issues Continuously improve platform performance, security, and reliability Stay current with emerging technologies and drive innovation and adoption Communicate complex technical ideas to both technical and non-technical stakeholders If you possess a combination of some of the following skills, then LET'S TALK! Strong experience designing and managing cloud native platforms on Google Cloud Platform (GCP) Hands-on expertise with services such as: Compute Engine, GKE (Kubernetes), Cloud Storage VPC networking, IAM, Cloud Functions Proven experience with Infrastructure as Code (e.g., Terraform) Strong background in CI/CD pipeline design (e.g., Jenkins, GitLab CI, Cloud Build) Proficiency in scripting/programming (e.g., Python, Bash, Go) Experience with containerisation & orchestration (Docker, Kubernetes) Solid understanding of cloud networking, security, and identity management Experience with monitoring & observability tools (Prometheus, Grafana, ELK, or similar) Strong Git/version control experience Exposure to the following skills is advantageous but not essential: - Experience with multi-cloud or hybrid cloud environments (AWS, Azure, on-premise). Exposure to service mesh technologies (e.g., Istio, Linkerd) and configuration management tools (e.g., Chef, Puppet). Knowledge of site reliability engineering (SRE) principles and practices. Familiarity with GCP AI/ML services (AI Platform, AutoML). In return, you will be rewarded with ongoing career development and a market leading benefits package in a flexible, hybrid working environment. What you need to do now If you're interested in this role, click 'apply now' to forward an up-to-date copy of your CV, or call us now. If this job isn't quite right for you, but you are looking for a new position, please contact us for a confidential discussion about your career. Hays Specialist Recruitment Limited acts as an employment agency for permanent recruitment and employment business for the supply of temporary workers. By applying for this job you accept the T&C's, Privacy Policy and Disclaimers which can be found at (url removed)
14/07/2026
Full time
Prestigious opportunity for a talented and experienced Platform Engineer to join a rapidly growing digital engineering team delivering cutting edge solutions across a diverse portfolio of clients.This is hands on applying your deep engineering and architectural expertise to design, build, and evolve scalable platform solutions that drive digital transformation within some of the UK's most exciting organisations.Collaborating closely with highly skilled, cross functional teams, you will work alongside engineers, architects, and product specialists to push boundaries and deliver impactful, high quality outcomes. Key responsibilities: Design, build, and maintain robust cloud platforms and CI/CD pipelines Contribute across the full software development lifecycle Provide technical leadership and input into system and architecture design Collaborate with cross functional teams to deliver scalable, reliable solutions. Write clean, well documented code and contribute to technical documentation Proactively monitor, troubleshoot, and resolve production issues Continuously improve platform performance, security, and reliability Stay current with emerging technologies and drive innovation and adoption Communicate complex technical ideas to both technical and non-technical stakeholders If you possess a combination of some of the following skills, then LET'S TALK! Strong experience designing and managing cloud native platforms on Google Cloud Platform (GCP) Hands-on expertise with services such as: Compute Engine, GKE (Kubernetes), Cloud Storage VPC networking, IAM, Cloud Functions Proven experience with Infrastructure as Code (e.g., Terraform) Strong background in CI/CD pipeline design (e.g., Jenkins, GitLab CI, Cloud Build) Proficiency in scripting/programming (e.g., Python, Bash, Go) Experience with containerisation & orchestration (Docker, Kubernetes) Solid understanding of cloud networking, security, and identity management Experience with monitoring & observability tools (Prometheus, Grafana, ELK, or similar) Strong Git/version control experience Exposure to the following skills is advantageous but not essential: - Experience with multi-cloud or hybrid cloud environments (AWS, Azure, on-premise). Exposure to service mesh technologies (e.g., Istio, Linkerd) and configuration management tools (e.g., Chef, Puppet). Knowledge of site reliability engineering (SRE) principles and practices. Familiarity with GCP AI/ML services (AI Platform, AutoML). In return, you will be rewarded with ongoing career development and a market leading benefits package in a flexible, hybrid working environment. What you need to do now If you're interested in this role, click 'apply now' to forward an up-to-date copy of your CV, or call us now. If this job isn't quite right for you, but you are looking for a new position, please contact us for a confidential discussion about your career. Hays Specialist Recruitment Limited acts as an employment agency for permanent recruitment and employment business for the supply of temporary workers. By applying for this job you accept the T&C's, Privacy Policy and Disclaimers which can be found at (url removed)
Government Digital & Data
SRE Squad Lead - Department for Business and Trade - G7
Government Digital & Data
Location Belfast, Birmingham, Cardiff, Darlington, Edinburgh, London, Salford About the job Job summary The Department for Business and Trade (DBT) has a clear mission - to grow the economy. Our role is to help businesses invest, grow and export to create jobs and opportunities right across the country. We do this in three ways. Firstly, we help to build a strong, competitive business environment, where consumers are protected and companies rewarded for treating their employees properly. Secondly, we open international markets and ensure resilient supply chains. This can be through Free Trade Agreements, trade facilitation and multilateral agreements. Finally, we work in partnership with businesses every day, providing advance, finance and deal-making support to those looking to start up, invest, export and grow. The Digital, Data and Technology (DDaT) directorate develops and operates tools and services to support us in this mission. About the role As a Senior Site Reliability Engineer Manager, you will lead the design, delivery and continuous improvement of reliable, scalable and secure platform services that underpin critical DBT digital products. Working closely with multidisciplinary, agile teams, you'll ensure development teams have the tools and support they need from observability and monitoring through to CI/CD pipelines, so services are resilient, performant and centred around user needs. You'll champion good engineering practices, helping teams adopt service-level thinking using metrics, service-level indicators (SLIs), objectives (SLOs) and error budgets to support informed, collaborative decision making. This is a people-focused leadership role where you will create an environment in which engineers can do their best work. You will line manage and develop a team of Site Reliability Engineers, supporting their growth and wellbeing, while also acting as a senior technical leader across the wider DDaT community. You'll work in partnership with product managers, architects and delivery colleagues to shape platform strategy, improve reliability and reduce operational burden. Alongside hands-on engineering, you will help build and scale our global platform, support live services through an on-call rota, and lead improvements such as enhancing observability and streamlining deployment processes to improve service quality and delivery outcomes. Job description You will: Lead and support a team of Site Reliability Engineers, setting clear direction while fostering an inclusive, collaborative and high-performing team culture. Build strong working relationships with product, delivery and architecture colleagues to ensure platform services meet business and user needs. Provide technical leadership across DevOps/SRE practices, guiding teams to adopt approaches that support reliability, sustainability and continuous improvement. Coach, mentor and support engineers across DDaT, contributing to a supportive and diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets in a pragmatic and user-focused way. Support the development and continual improvement of CI/CD pipelines to enable safe, frequent and low-risk delivery of changes. Oversee live service reliability, supporting teams through incident and problem management while encouraging a learning-focused, blameless culture. Ensure security, resilience and compliance considerations are understood and embedded into engineering practices. What tech will you be using? AWS and Azure GitHub Actions, AWS CodePipelines/CodeBuild Terraform Docker, Elastic Container Service (ECS) and Elastic Container Registry (ECR) ElasticSearch/OpenSearch Python and Django framework PostgreSQL as a service (Amazon RDS) Datadog, Logstash Redis/Elasticache Person specification It is essential that you have: Experience leading, supporting and developing engineers, including line management or strong mentoring experience. Strong communication skills, with the ability to explain technical concepts clearly and build effective relationships with a range of stakeholders. Experience of working with cloud platforms such as AWS, Azure or Google Cloud, applying modern DevOps and SRE practices. Experience designing and delivering infrastructure-as-code solutions using tools such as Terraform, CloudFormation or similar. Ability to write clean, maintainable and well-tested code in at least one programming language. Experience designing, operating and improving distributed systems, with a focus on reliability, performance and user impact.
14/07/2026
Full time
Location Belfast, Birmingham, Cardiff, Darlington, Edinburgh, London, Salford About the job Job summary The Department for Business and Trade (DBT) has a clear mission - to grow the economy. Our role is to help businesses invest, grow and export to create jobs and opportunities right across the country. We do this in three ways. Firstly, we help to build a strong, competitive business environment, where consumers are protected and companies rewarded for treating their employees properly. Secondly, we open international markets and ensure resilient supply chains. This can be through Free Trade Agreements, trade facilitation and multilateral agreements. Finally, we work in partnership with businesses every day, providing advance, finance and deal-making support to those looking to start up, invest, export and grow. The Digital, Data and Technology (DDaT) directorate develops and operates tools and services to support us in this mission. About the role As a Senior Site Reliability Engineer Manager, you will lead the design, delivery and continuous improvement of reliable, scalable and secure platform services that underpin critical DBT digital products. Working closely with multidisciplinary, agile teams, you'll ensure development teams have the tools and support they need from observability and monitoring through to CI/CD pipelines, so services are resilient, performant and centred around user needs. You'll champion good engineering practices, helping teams adopt service-level thinking using metrics, service-level indicators (SLIs), objectives (SLOs) and error budgets to support informed, collaborative decision making. This is a people-focused leadership role where you will create an environment in which engineers can do their best work. You will line manage and develop a team of Site Reliability Engineers, supporting their growth and wellbeing, while also acting as a senior technical leader across the wider DDaT community. You'll work in partnership with product managers, architects and delivery colleagues to shape platform strategy, improve reliability and reduce operational burden. Alongside hands-on engineering, you will help build and scale our global platform, support live services through an on-call rota, and lead improvements such as enhancing observability and streamlining deployment processes to improve service quality and delivery outcomes. Job description You will: Lead and support a team of Site Reliability Engineers, setting clear direction while fostering an inclusive, collaborative and high-performing team culture. Build strong working relationships with product, delivery and architecture colleagues to ensure platform services meet business and user needs. Provide technical leadership across DevOps/SRE practices, guiding teams to adopt approaches that support reliability, sustainability and continuous improvement. Coach, mentor and support engineers across DDaT, contributing to a supportive and diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets in a pragmatic and user-focused way. Support the development and continual improvement of CI/CD pipelines to enable safe, frequent and low-risk delivery of changes. Oversee live service reliability, supporting teams through incident and problem management while encouraging a learning-focused, blameless culture. Ensure security, resilience and compliance considerations are understood and embedded into engineering practices. What tech will you be using? AWS and Azure GitHub Actions, AWS CodePipelines/CodeBuild Terraform Docker, Elastic Container Service (ECS) and Elastic Container Registry (ECR) ElasticSearch/OpenSearch Python and Django framework PostgreSQL as a service (Amazon RDS) Datadog, Logstash Redis/Elasticache Person specification It is essential that you have: Experience leading, supporting and developing engineers, including line management or strong mentoring experience. Strong communication skills, with the ability to explain technical concepts clearly and build effective relationships with a range of stakeholders. Experience of working with cloud platforms such as AWS, Azure or Google Cloud, applying modern DevOps and SRE practices. Experience designing and delivering infrastructure-as-code solutions using tools such as Terraform, CloudFormation or similar. Ability to write clean, maintainable and well-tested code in at least one programming language. Experience designing, operating and improving distributed systems, with a focus on reliability, performance and user impact.
Lead Product Manager AIOPs
SwiftCruit
# Lead Product Manager AIOPsOnsite Lead Level Full TimePosted: Yesterday SkillsDistributed SystemsReliability EngineeringAlgorithmsHTTPAWSAzureGCPLangChainRAGPrompt EngineeringMachine LearningLLMGenerative AIDatadogNew RelicSplunkServiceNow Company LocationLondon, UK, United Kingdom Remote Work PolicyOnsite# About the Role: Grade Level (for internal use): 11 The Team: The AIOps team is responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights.DTS Platform & Tools - Service Enablement: We serve as thought leaders in AIOps, partnering across IT Operations, SRE, engineering, infrastructure, service management, and application teams to solve enterprise operational challenges. Our mission is to improve reliability, reduce operational complexity, optimize technology investments, and enable more proactive and resilient technology operations by applying AI. Responsibilities and Impact: Own and execute the AIOps product roadmap, aligning priorities with enterprise reliability goals, operational maturity, adoption targets, and measurable business outcomes. Translate operational needs into clear product requirements, use cases, user stories, acceptance criteria, and prioritized backlog items across the AIOps product lifecycle. Ownership of high-impact AIOps capabilities, including noise reduction, event management, correlation, anomaly detection, root cause analysis, predictive alerting, and self-healing remediation. Partner across Operations, infrastructure, SRE, platform engineering, service management, application teams, security, and vendors to operationalize AIOps capabilities and drive adoption at enterprise scale. Define and track success metrics for alert quality, incident reduction, MTTD, MTTR, service health, automation coverage, platform adoption, and operational efficiency. Present roadmap progress, risks, decisions, and outcome metrics to senior stakeholders to influence improvements to tooling, workflow, and the operating model. Analytical and data-driven mindset with strong problem-solving, prioritization, and decision-making skills. What We're Looking For: Basic Required Qualifications: 10+ years of experience in product management, IT operations, SRE, observability, platform engineering, or related enterprise technology roles. Strong understanding of AIOps concepts, including event correlation, anomaly detection, root cause analysis, noise reduction, predictive analytics, and automated remediation. Experience defining product roadmaps, managing requirements and backlog priorities, and delivering technical platform capabilities across the product lifecycle. Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent practical experience. Machine Learning: Expertise in time-series forecasting, clustering, and correlation algorithms. Generative AI: Experience with LLM prompting, Retrieval-Augmented Generation (RAG), and LangChain. Cloud Infrastructure: Deep understanding of AWS, Azure, or GCP infrastructure components Additional Preferred Qualifications: Experience with enterprise observability and AIOps platforms such as ServiceNow ITOM, Splunk, Dynatrace, Datadog, BigPanda, New Relic, or similar platforms. Familiarity with cloud-native platforms, Container orchestration systems, distributed systems, service mapping, and telemetry standards. Experience defining KPIs, SLOs, product adoption measures, and operational maturity frameworks for enterprise reliability and service performance. Relevant certifications in product management, cloud, ITSM, observability, Agile, or platform engineering are a plus. What's In It For You? Our Mission: Advancing Essential Intelligence. Our People: We're more than 35,000 strong worldwide-so we're able to understand nuances while having a broad perspective. Our team is driven by curiosity and a shared belief that Essential Intelligence can help build a more prosperous future for us all.From finding new ways to measure sustainability to analyzing energy transition across the supply chain to building workflow solutions that make it easy to tap into insight and apply it. We are changing the way people see things and empowering them to make an impact on the world we live in. We're committed to a more equitable future and to helping our customers find new, sustainable ways of doing business. Join us and help create the critical insights that truly make a difference. Our Values: Integrity, Discovery, Partnership Throughout our history, the world's leading organizations have relied on us for the Essential Intelligence they need to make confident decisions about the road ahead. We start with a foundation of integrity in all we do, bring a spirit of discovery to our work, and collaborate in close partnership with each other and our customers to achieve shared goals. Benefits: We take care of you, so you can take care of business. We care about our people. That's why we provide everything you-and your career-need to thrive at S&P Global. Our benefits include: Health & Wellness: Health care coverage designed for the mind and body. Flexible Downtime: Generous time off helps keep you energized for your time on. Continuous Learning: Access a wealth of resources to grow your career and learn valuable new skills. Invest in Your Future: Secure your financial future through competitive pay, retirement planning, a continuing education program with a company-matched student loan contribution, and financial wellness programs. Family Friendly Perks: It's not just about you. S&P Global has perks for your partners and little ones, too, with some best-in class benefits for families. Beyond the Basics: From retail discounts to referral incentive awards-small perks can make a big difference.For more information on benefits by country visit: Global Hiring and Opportunity at S&P Global: At S&P Global, we are committed to fostering a connected and engaged workplace where all individuals have access to opportunities based on their skills, experience, and contributions. Our hiring practices emphasize fairness, transparency, and merit, ensuring that we attract and retain top talent. By valuing different perspectives and promoting a culture of respect and collaboration, we drive innovation and power global markets. Recruitment Fraud Alert: If you receive an email from a domain or any other regionally based domains, it is a scam and should be reported to . S&P Global never requires any candidate to pay money for job applications, interviews, offer letters, "pre-employment training" or for equipment/delivery of equipment. Stay informed and protect yourself from recruitment fraud by reviewing our guidelines, fraudulent domains, and how to report suspicious activity here. - Equal Opportunity Employer S&P Global is an equal opportunity employer and all qualified candidates will receive consideration for employment without regard to race/ethnicity, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, marital status, military veteran status, unemployment status, or any other status protected by law. Only electronic job submissions will be considered for employment.If you need an accommodation during the application process due to a disability, please send an email to: and your request will be forwarded to the appropriate person.
14/07/2026
Full time
# Lead Product Manager AIOPsOnsite Lead Level Full TimePosted: Yesterday SkillsDistributed SystemsReliability EngineeringAlgorithmsHTTPAWSAzureGCPLangChainRAGPrompt EngineeringMachine LearningLLMGenerative AIDatadogNew RelicSplunkServiceNow Company LocationLondon, UK, United Kingdom Remote Work PolicyOnsite# About the Role: Grade Level (for internal use): 11 The Team: The AIOps team is responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights.DTS Platform & Tools - Service Enablement: We serve as thought leaders in AIOps, partnering across IT Operations, SRE, engineering, infrastructure, service management, and application teams to solve enterprise operational challenges. Our mission is to improve reliability, reduce operational complexity, optimize technology investments, and enable more proactive and resilient technology operations by applying AI. Responsibilities and Impact: Own and execute the AIOps product roadmap, aligning priorities with enterprise reliability goals, operational maturity, adoption targets, and measurable business outcomes. Translate operational needs into clear product requirements, use cases, user stories, acceptance criteria, and prioritized backlog items across the AIOps product lifecycle. Ownership of high-impact AIOps capabilities, including noise reduction, event management, correlation, anomaly detection, root cause analysis, predictive alerting, and self-healing remediation. Partner across Operations, infrastructure, SRE, platform engineering, service management, application teams, security, and vendors to operationalize AIOps capabilities and drive adoption at enterprise scale. Define and track success metrics for alert quality, incident reduction, MTTD, MTTR, service health, automation coverage, platform adoption, and operational efficiency. Present roadmap progress, risks, decisions, and outcome metrics to senior stakeholders to influence improvements to tooling, workflow, and the operating model. Analytical and data-driven mindset with strong problem-solving, prioritization, and decision-making skills. What We're Looking For: Basic Required Qualifications: 10+ years of experience in product management, IT operations, SRE, observability, platform engineering, or related enterprise technology roles. Strong understanding of AIOps concepts, including event correlation, anomaly detection, root cause analysis, noise reduction, predictive analytics, and automated remediation. Experience defining product roadmaps, managing requirements and backlog priorities, and delivering technical platform capabilities across the product lifecycle. Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent practical experience. Machine Learning: Expertise in time-series forecasting, clustering, and correlation algorithms. Generative AI: Experience with LLM prompting, Retrieval-Augmented Generation (RAG), and LangChain. Cloud Infrastructure: Deep understanding of AWS, Azure, or GCP infrastructure components Additional Preferred Qualifications: Experience with enterprise observability and AIOps platforms such as ServiceNow ITOM, Splunk, Dynatrace, Datadog, BigPanda, New Relic, or similar platforms. Familiarity with cloud-native platforms, Container orchestration systems, distributed systems, service mapping, and telemetry standards. Experience defining KPIs, SLOs, product adoption measures, and operational maturity frameworks for enterprise reliability and service performance. Relevant certifications in product management, cloud, ITSM, observability, Agile, or platform engineering are a plus. What's In It For You? Our Mission: Advancing Essential Intelligence. Our People: We're more than 35,000 strong worldwide-so we're able to understand nuances while having a broad perspective. Our team is driven by curiosity and a shared belief that Essential Intelligence can help build a more prosperous future for us all.From finding new ways to measure sustainability to analyzing energy transition across the supply chain to building workflow solutions that make it easy to tap into insight and apply it. We are changing the way people see things and empowering them to make an impact on the world we live in. We're committed to a more equitable future and to helping our customers find new, sustainable ways of doing business. Join us and help create the critical insights that truly make a difference. Our Values: Integrity, Discovery, Partnership Throughout our history, the world's leading organizations have relied on us for the Essential Intelligence they need to make confident decisions about the road ahead. We start with a foundation of integrity in all we do, bring a spirit of discovery to our work, and collaborate in close partnership with each other and our customers to achieve shared goals. Benefits: We take care of you, so you can take care of business. We care about our people. That's why we provide everything you-and your career-need to thrive at S&P Global. Our benefits include: Health & Wellness: Health care coverage designed for the mind and body. Flexible Downtime: Generous time off helps keep you energized for your time on. Continuous Learning: Access a wealth of resources to grow your career and learn valuable new skills. Invest in Your Future: Secure your financial future through competitive pay, retirement planning, a continuing education program with a company-matched student loan contribution, and financial wellness programs. Family Friendly Perks: It's not just about you. S&P Global has perks for your partners and little ones, too, with some best-in class benefits for families. Beyond the Basics: From retail discounts to referral incentive awards-small perks can make a big difference.For more information on benefits by country visit: Global Hiring and Opportunity at S&P Global: At S&P Global, we are committed to fostering a connected and engaged workplace where all individuals have access to opportunities based on their skills, experience, and contributions. Our hiring practices emphasize fairness, transparency, and merit, ensuring that we attract and retain top talent. By valuing different perspectives and promoting a culture of respect and collaboration, we drive innovation and power global markets. Recruitment Fraud Alert: If you receive an email from a domain or any other regionally based domains, it is a scam and should be reported to . S&P Global never requires any candidate to pay money for job applications, interviews, offer letters, "pre-employment training" or for equipment/delivery of equipment. Stay informed and protect yourself from recruitment fraud by reviewing our guidelines, fraudulent domains, and how to report suspicious activity here. - Equal Opportunity Employer S&P Global is an equal opportunity employer and all qualified candidates will receive consideration for employment without regard to race/ethnicity, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, marital status, military veteran status, unemployment status, or any other status protected by law. Only electronic job submissions will be considered for employment.If you need an accommodation during the application process due to a disability, please send an email to: and your request will be forwarded to the appropriate person.
Lead Product Manager AIOPs
S&P Global, Inc.
About the Role: Grade Level (for internal use): 11 The Team: The AIOps team is responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights. DTS Platform & Tools - Service Enablement: We serve as thought leaders in AIOps, partnering across IT Operations, SRE, engineering, infrastructure, service management, and application teams to solve enterprise operational challenges. Our mission is to improve reliability, reduce operational complexity, optimize technology investments, and enable more proactive and resilient technology operations by applying AI. Responsibilities and Impact: Own and execute the AIOps product roadmap, aligning priorities with enterprise reliability goals, operational maturity, adoption targets, and measurable business outcomes. Translate operational needs into clear product requirements, use cases, user stories, acceptance criteria, and prioritized backlog items across the AIOps product lifecycle. Ownership of high-impact AIOps capabilities, including noise reduction, event management, correlation, anomaly detection, root cause analysis, predictive alerting, and self-healing remediation. Partner across Operations, infrastructure, SRE, platform engineering, service management, application teams, security, and vendors to operationalize AIOps capabilities and drive adoption at enterprise scale. Define and track success metrics for alert quality, incident reduction, MTTD, MTTR, service health, automation coverage, platform adoption, and operational efficiency. Present roadmap progress, risks, decisions, and outcome metrics to senior stakeholders to influence improvements to tooling, workflow, and the operating model. Analytical and data driven mindset with strong problem solving, prioritization, and decision making skills. What We're Looking For: Basic Required Qualifications: 10+ years of experience in product management, IT operations, SRE, observability, platform engineering, or related enterprise technology roles. Strong understanding of AIOps concepts, including event correlation, anomaly detection, root cause analysis, noise reduction, predictive analytics, and automated remediation. Experience defining product roadmaps, managing requirements and backlog priorities, and delivering technical platform capabilities across the product lifecycle. Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent practical experience. Machine Learning: Expertise in time series forecasting, clustering, and correlation algorithms. Generative AI: Experience with LLM prompting, Retrieval Augmented Generation (RAG), and LangChain. Cloud Infrastructure: Deep understanding of AWS, Azure, or GCP infrastructure components Additional Preferred Qualifications: Experience with enterprise observability and AIOps platforms such as ServiceNow ITOM, Splunk, Dynatrace, Datadog, BigPanda, New Relic, or similar platforms. Familiarity with cloud native platforms, Container orchestration systems, distributed systems, service mapping, and telemetry standards. Experience defining KPIs, SLOs, product adoption measures, and operational maturity frameworks for enterprise reliability and service performance. Relevant certifications in product management, cloud, ITSM, observability, Agile, or platform engineering are a plus. Benefits: Health & Wellness: Health care coverage designed for the mind and body. Flexible Downtime: Generous time off helps keep you energized for your time on. Continuous Learning: Access a wealth of resources to grow your career and learn valuable new skills. Invest in Your Future: Secure your financial future through competitive pay, retirement planning, a continuing education program with a company matched student loan contribution, and financial wellness programs. Family Friendly Perks: It's not just about you. S&P Global has perks for your partners and little ones, too, with some best in class benefits for families. Beyond the Basics: From retail discounts to referral incentive awards-small perks can make a big difference. For more information on benefits by country visit: Equal Opportunity Employer S&P Global is an equal opportunity employer and all qualified candidates will receive consideration for employment without regard to race/ethnicity, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, marital status, military veteran status, unemployment status, or any other status protected by law. Only electronic job submissions will be considered for employment. If you need an accommodation during the application process due to a disability, please send an email to: and your request will be forwarded to the appropriate person.
14/07/2026
Full time
About the Role: Grade Level (for internal use): 11 The Team: The AIOps team is responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights. DTS Platform & Tools - Service Enablement: We serve as thought leaders in AIOps, partnering across IT Operations, SRE, engineering, infrastructure, service management, and application teams to solve enterprise operational challenges. Our mission is to improve reliability, reduce operational complexity, optimize technology investments, and enable more proactive and resilient technology operations by applying AI. Responsibilities and Impact: Own and execute the AIOps product roadmap, aligning priorities with enterprise reliability goals, operational maturity, adoption targets, and measurable business outcomes. Translate operational needs into clear product requirements, use cases, user stories, acceptance criteria, and prioritized backlog items across the AIOps product lifecycle. Ownership of high-impact AIOps capabilities, including noise reduction, event management, correlation, anomaly detection, root cause analysis, predictive alerting, and self-healing remediation. Partner across Operations, infrastructure, SRE, platform engineering, service management, application teams, security, and vendors to operationalize AIOps capabilities and drive adoption at enterprise scale. Define and track success metrics for alert quality, incident reduction, MTTD, MTTR, service health, automation coverage, platform adoption, and operational efficiency. Present roadmap progress, risks, decisions, and outcome metrics to senior stakeholders to influence improvements to tooling, workflow, and the operating model. Analytical and data driven mindset with strong problem solving, prioritization, and decision making skills. What We're Looking For: Basic Required Qualifications: 10+ years of experience in product management, IT operations, SRE, observability, platform engineering, or related enterprise technology roles. Strong understanding of AIOps concepts, including event correlation, anomaly detection, root cause analysis, noise reduction, predictive analytics, and automated remediation. Experience defining product roadmaps, managing requirements and backlog priorities, and delivering technical platform capabilities across the product lifecycle. Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent practical experience. Machine Learning: Expertise in time series forecasting, clustering, and correlation algorithms. Generative AI: Experience with LLM prompting, Retrieval Augmented Generation (RAG), and LangChain. Cloud Infrastructure: Deep understanding of AWS, Azure, or GCP infrastructure components Additional Preferred Qualifications: Experience with enterprise observability and AIOps platforms such as ServiceNow ITOM, Splunk, Dynatrace, Datadog, BigPanda, New Relic, or similar platforms. Familiarity with cloud native platforms, Container orchestration systems, distributed systems, service mapping, and telemetry standards. Experience defining KPIs, SLOs, product adoption measures, and operational maturity frameworks for enterprise reliability and service performance. Relevant certifications in product management, cloud, ITSM, observability, Agile, or platform engineering are a plus. Benefits: Health & Wellness: Health care coverage designed for the mind and body. Flexible Downtime: Generous time off helps keep you energized for your time on. Continuous Learning: Access a wealth of resources to grow your career and learn valuable new skills. Invest in Your Future: Secure your financial future through competitive pay, retirement planning, a continuing education program with a company matched student loan contribution, and financial wellness programs. Family Friendly Perks: It's not just about you. S&P Global has perks for your partners and little ones, too, with some best in class benefits for families. Beyond the Basics: From retail discounts to referral incentive awards-small perks can make a big difference. For more information on benefits by country visit: Equal Opportunity Employer S&P Global is an equal opportunity employer and all qualified candidates will receive consideration for employment without regard to race/ethnicity, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, marital status, military veteran status, unemployment status, or any other status protected by law. Only electronic job submissions will be considered for employment. If you need an accommodation during the application process due to a disability, please send an email to: and your request will be forwarded to the appropriate person.
Hays Specialist Recruitment
Platform Engineer (GCP)
Hays Specialist Recruitment Manchester, Lancashire
Prestigious opportunity for a talented and experienced Platform Engineer to join a rapidly growing digital engineering team delivering cutting edge solutions across a diverse portfolio of clients.This is hands on applying your deep engineering and architectural expertise to design, build, and evolve scalable platform solutions that drive digital transformation within some of the UK's most exciting organisations.Collaborating closely with highly skilled, cross functional teams, you will work alongside engineers, architects, and product specialists to push boundaries and deliver impactful, high quality outcomes. Key responsibilities: Design, build, and maintain robust cloud platforms and CI/CD pipelines Contribute across the full software development life cycle Provide technical leadership and input into system and architecture design Collaborate with cross functional teams to deliver scalable, reliable solutions. Write clean, well documented code and contribute to technical documentation Proactively monitor, troubleshoot, and resolve production issues Continuously improve platform performance, security, and reliability Stay current with emerging technologies and drive innovation and adoption Communicate complex technical ideas to both technical and non-technical stakeholders If you possess a combination of some of the following skills, then LET'S TALK! Strong experience designing and managing cloud native platforms on Google Cloud Platform (GCP) Hands-on expertise with services such as: Compute Engine, GKE (Kubernetes), Cloud Storage VPC networking, IAM, Cloud Functions Proven experience with Infrastructure as Code (eg, Terraform) Strong background in CI/CD pipeline design (eg, Jenkins, GitLab CI, Cloud Build) Proficiency in Scripting/programming (eg, Python, Bash, Go) Experience with containerisation & orchestration (Docker, Kubernetes) Solid understanding of cloud networking, security, and identity management Experience with monitoring & observability tools (Prometheus, Grafana, ELK, or similar) Strong Git/version control experience Exposure to the following skills is advantageous but not essential: - Experience with multi-cloud or hybrid cloud environments (AWS, Azure, on-premise). Exposure to service mesh technologies (eg, Istio, Linkerd) and configuration management tools (eg, Chef, Puppet). Knowledge of site reliability engineering (SRE) principles and practices. Familiarity with GCP AI/ML services (AI Platform, AutoML). In return, you will be rewarded with ongoing career development and a market leading benefits package in a flexible, hybrid working environment. What you need to do now If you're interested in this role, click 'apply now' to forward an up-to-date copy of your CV, or call us now. Hays Specialist Recruitment Limited acts as an employment agency for permanent recruitment and employment business for the supply of temporary workers. By applying for this job you accept the T&C's, Privacy Policy and Disclaimers which can be found on our website.
14/07/2026
Full time
Prestigious opportunity for a talented and experienced Platform Engineer to join a rapidly growing digital engineering team delivering cutting edge solutions across a diverse portfolio of clients.This is hands on applying your deep engineering and architectural expertise to design, build, and evolve scalable platform solutions that drive digital transformation within some of the UK's most exciting organisations.Collaborating closely with highly skilled, cross functional teams, you will work alongside engineers, architects, and product specialists to push boundaries and deliver impactful, high quality outcomes. Key responsibilities: Design, build, and maintain robust cloud platforms and CI/CD pipelines Contribute across the full software development life cycle Provide technical leadership and input into system and architecture design Collaborate with cross functional teams to deliver scalable, reliable solutions. Write clean, well documented code and contribute to technical documentation Proactively monitor, troubleshoot, and resolve production issues Continuously improve platform performance, security, and reliability Stay current with emerging technologies and drive innovation and adoption Communicate complex technical ideas to both technical and non-technical stakeholders If you possess a combination of some of the following skills, then LET'S TALK! Strong experience designing and managing cloud native platforms on Google Cloud Platform (GCP) Hands-on expertise with services such as: Compute Engine, GKE (Kubernetes), Cloud Storage VPC networking, IAM, Cloud Functions Proven experience with Infrastructure as Code (eg, Terraform) Strong background in CI/CD pipeline design (eg, Jenkins, GitLab CI, Cloud Build) Proficiency in Scripting/programming (eg, Python, Bash, Go) Experience with containerisation & orchestration (Docker, Kubernetes) Solid understanding of cloud networking, security, and identity management Experience with monitoring & observability tools (Prometheus, Grafana, ELK, or similar) Strong Git/version control experience Exposure to the following skills is advantageous but not essential: - Experience with multi-cloud or hybrid cloud environments (AWS, Azure, on-premise). Exposure to service mesh technologies (eg, Istio, Linkerd) and configuration management tools (eg, Chef, Puppet). Knowledge of site reliability engineering (SRE) principles and practices. Familiarity with GCP AI/ML services (AI Platform, AutoML). In return, you will be rewarded with ongoing career development and a market leading benefits package in a flexible, hybrid working environment. What you need to do now If you're interested in this role, click 'apply now' to forward an up-to-date copy of your CV, or call us now. Hays Specialist Recruitment Limited acts as an employment agency for permanent recruitment and employment business for the supply of temporary workers. By applying for this job you accept the T&C's, Privacy Policy and Disclaimers which can be found on our website.
London Stock Exchange Group
Senior Site Reliability Engineer
London Stock Exchange Group Nottingham, Nottinghamshire
Role Profile: We are evolving our Site Reliability Engineering capabilities to strengthen reliability, observability, security, and operational excellence across our Risk Intelligence division.As a Senior SRE , you will be a senior hands on technical person help shape the foundations of reliability across both new and existing platforms. You will collaborate with Architecture, Engineering, Security, and Platform teams to ensure reliability is built into systems from day one.While this is not a people management you will work closely with global teams and may occasionally be called upon for major incidents or critical issues. This position requires a highly proactive, hard-working expert with strong leadership presence and ownership of platform reliability outcomes. Key Responsibilities We are looking for a person who is passionate about reliability engineering and who bring a continuous improvement approach to everything they do!Lead the establishment of SRE foundations for new projects building environments, monitoring, alerting, and ensuring operational readiness from day one.Define, implement, and champion observability standards, tooling, and guidelines across metrics, logs, traces, and SLIs/SLOs.Design and evolve monitoring and alerting solutions that improve visibility, reduce toil, and strengthen system health.Continuously drive reliability improvements across our environments through incident reduction, performance tuning, and building resilient patterns.Partner with Security teams to ensure our platforms meet compliance, security, and risk management expectations.Influence architectural and design decisions through data driven cloud cost optimization and efficiency initiatives.Be a technical leader and mentor supporting engineers, shaping engineering standards, and fostering a culture of learning and development. PERSON SPECIFICATION Education: Bachelor's Degree or equivalent experience in Computer Science, Engineering, or a related field Required skills and experience 5+ years of hands-on technical experience in SRE, Platform Engineering, Infrastructure, or related roles Strong experience with AWS (or Azure) , including services such as EKS, ECS, EC2, networking, IAM, and managed services Solid understanding of cloud security principles and experience collaborating with security teams. Strong background in Linux systems administrations. Proven experience designing and operating observability platforms , including monitoring, logging, and alerting Hands-on experience with Datadog for metrics, logs, APM, and alerting Strong understanding of SRE principles , including SLOs, error budgets, incident management, and reliability engineering Experience working closely with architecture and engineering teams on system design and delivery Experience with cloud cost optimization strategies and tooling Good to have Skills Experience supporting multi-cloud or hybrid environments Exposure to Infrastructure as Code (e.g., Terraform, CloudFormation) Experience in large-scale, complex, or regulated environments Knowledge of vector databases and RAG architectures for building internal SRE knowledge assistants. Knowledge of Generative AI and LLM platforms (e.g., Claude, Amazon Bedrock) Person Specification Strong technical authority with the ability to influence design and operational decisions Highly collaborative, comfortable working across architecture, engineering, security, and operations teams Calm and methodical under pressure, especially during incidents and critical issues Pragmatic problem-solver who balances reliability, security, cost, and delivery speed Clear communicator, able to explain complex technical concepts to diverse audiences Career Stage: Senior Associate London Stock Exchange Group (LSEG) Information: Join us and be part of a team that values innovation, quality, and continuous improvement. If you're ready to take your career to the next level and make a significant impact, we'd love to hear from you.LSEG is a leading global financial markets infrastructure and data provider. Our purpose is driving financial stability, empowering economies and enabling customers to create sustainable growth.Our purpose is the foundation on which our culture is built. Our values of Integrity, Partnership , Excellence and Change underpin our purpose and set the standard for everything we do, every day. They go to the heart of who we are and guide our decision making and everyday actions.Working with us means that you will be part of a dynamic organisation of 25,000 people across 65 countries. However, we will value your individuality and enable you to bring your true self to work so you can help enrich our diverse workforce.We are proud to be an equal opportunities employer. This means that we do not discriminate on the basis of anyone's race, religion, colour, national origin, gender, sexual orientation, gender identity, gender expression, age, marital status, veteran status, pregnancy or disability, or any other basis protected under applicable law. Conforming with applicable law, we can reasonably accommodate applicants' and employees' religious practices and beliefs, as well as mental health or physical disability needs.You will be part of a collaborative and creative culture where we encourage new ideas. We are committed to sustainability across our global business and we are proud to partner with our customers to help them meet their sustainability objectives. Our charity, the LSEG Foundation provides charitable grants to community groups that help people access economic opportunities and build a secure future with financial independence. Colleagues can get involved through fundraising and volunteering.LSEG offers a range of tailored benefits and support, including healthcare, retirement planning, paid volunteering days and wellbeing initiatives.Please take a moment to read this carefully, as it describes what personal information London Stock Exchange Group (LSEG) (we) may hold about you, what it's used for, and how it's obtained, .If you are submitting as a Recruitment Agency Partner, it is essential and your responsibility to ensure that candidates applying to LSEG are aware of this privacy notice.LSEG (London Stock Exchange Group) is a leading global financial markets infrastructure and data provider. Our purpose is driving financial stability, empowering economies and enabling customers to create sustainable growth. Our culture of connecting, creating opportunity and delivering excellence shapes how we think, how we do things and how we help our people fulfil their potential.
13/07/2026
Full time
Role Profile: We are evolving our Site Reliability Engineering capabilities to strengthen reliability, observability, security, and operational excellence across our Risk Intelligence division.As a Senior SRE , you will be a senior hands on technical person help shape the foundations of reliability across both new and existing platforms. You will collaborate with Architecture, Engineering, Security, and Platform teams to ensure reliability is built into systems from day one.While this is not a people management you will work closely with global teams and may occasionally be called upon for major incidents or critical issues. This position requires a highly proactive, hard-working expert with strong leadership presence and ownership of platform reliability outcomes. Key Responsibilities We are looking for a person who is passionate about reliability engineering and who bring a continuous improvement approach to everything they do!Lead the establishment of SRE foundations for new projects building environments, monitoring, alerting, and ensuring operational readiness from day one.Define, implement, and champion observability standards, tooling, and guidelines across metrics, logs, traces, and SLIs/SLOs.Design and evolve monitoring and alerting solutions that improve visibility, reduce toil, and strengthen system health.Continuously drive reliability improvements across our environments through incident reduction, performance tuning, and building resilient patterns.Partner with Security teams to ensure our platforms meet compliance, security, and risk management expectations.Influence architectural and design decisions through data driven cloud cost optimization and efficiency initiatives.Be a technical leader and mentor supporting engineers, shaping engineering standards, and fostering a culture of learning and development. PERSON SPECIFICATION Education: Bachelor's Degree or equivalent experience in Computer Science, Engineering, or a related field Required skills and experience 5+ years of hands-on technical experience in SRE, Platform Engineering, Infrastructure, or related roles Strong experience with AWS (or Azure) , including services such as EKS, ECS, EC2, networking, IAM, and managed services Solid understanding of cloud security principles and experience collaborating with security teams. Strong background in Linux systems administrations. Proven experience designing and operating observability platforms , including monitoring, logging, and alerting Hands-on experience with Datadog for metrics, logs, APM, and alerting Strong understanding of SRE principles , including SLOs, error budgets, incident management, and reliability engineering Experience working closely with architecture and engineering teams on system design and delivery Experience with cloud cost optimization strategies and tooling Good to have Skills Experience supporting multi-cloud or hybrid environments Exposure to Infrastructure as Code (e.g., Terraform, CloudFormation) Experience in large-scale, complex, or regulated environments Knowledge of vector databases and RAG architectures for building internal SRE knowledge assistants. Knowledge of Generative AI and LLM platforms (e.g., Claude, Amazon Bedrock) Person Specification Strong technical authority with the ability to influence design and operational decisions Highly collaborative, comfortable working across architecture, engineering, security, and operations teams Calm and methodical under pressure, especially during incidents and critical issues Pragmatic problem-solver who balances reliability, security, cost, and delivery speed Clear communicator, able to explain complex technical concepts to diverse audiences Career Stage: Senior Associate London Stock Exchange Group (LSEG) Information: Join us and be part of a team that values innovation, quality, and continuous improvement. If you're ready to take your career to the next level and make a significant impact, we'd love to hear from you.LSEG is a leading global financial markets infrastructure and data provider. Our purpose is driving financial stability, empowering economies and enabling customers to create sustainable growth.Our purpose is the foundation on which our culture is built. Our values of Integrity, Partnership , Excellence and Change underpin our purpose and set the standard for everything we do, every day. They go to the heart of who we are and guide our decision making and everyday actions.Working with us means that you will be part of a dynamic organisation of 25,000 people across 65 countries. However, we will value your individuality and enable you to bring your true self to work so you can help enrich our diverse workforce.We are proud to be an equal opportunities employer. This means that we do not discriminate on the basis of anyone's race, religion, colour, national origin, gender, sexual orientation, gender identity, gender expression, age, marital status, veteran status, pregnancy or disability, or any other basis protected under applicable law. Conforming with applicable law, we can reasonably accommodate applicants' and employees' religious practices and beliefs, as well as mental health or physical disability needs.You will be part of a collaborative and creative culture where we encourage new ideas. We are committed to sustainability across our global business and we are proud to partner with our customers to help them meet their sustainability objectives. Our charity, the LSEG Foundation provides charitable grants to community groups that help people access economic opportunities and build a secure future with financial independence. Colleagues can get involved through fundraising and volunteering.LSEG offers a range of tailored benefits and support, including healthcare, retirement planning, paid volunteering days and wellbeing initiatives.Please take a moment to read this carefully, as it describes what personal information London Stock Exchange Group (LSEG) (we) may hold about you, what it's used for, and how it's obtained, .If you are submitting as a Recruitment Agency Partner, it is essential and your responsibility to ensure that candidates applying to LSEG are aware of this privacy notice.LSEG (London Stock Exchange Group) is a leading global financial markets infrastructure and data provider. Our purpose is driving financial stability, empowering economies and enabling customers to create sustainable growth. Our culture of connecting, creating opportunity and delivering excellence shapes how we think, how we do things and how we help our people fulfil their potential.
Front Office Product Support
TP ICAP Group
Front Office Product Support page is loaded Front Office Product Supportlocations: Londontime type: Full timeposted on: Posted Yesterdayjob requisition id: R5215The TP ICAP Group is a world leading provider of market infrastructure.Our purpose is to provide clients with access to global financial and commodities markets, improving price discovery, liquidity, and distribution of data, through responsible and innovative solutions.Through our people and technology, we connect clients to superior liquidity and data solutions.The Group is home to a stable of premium brands. Collectively, TP ICAP is the largest interdealer broker in the world by revenue, the number one Energy & Commodities broker in the world, the world's leading provider of OTC data, and an award winning all-to-all trading platform.The Group operates from more than 60 offices in 27 countries. We are 5,300 people strong. We work as one to achieve our vision of being the world's most trusted, innovative, liquidity and data solutions specialist. Role Overview Liquidnet is looking for an Application Support engineer to work within the EMEA Front Office Support team. The team requires a motivated self-starter who has the technical skills to support a growing number of buy-side members utilising their FIX, Linux, Windows Server, DevOps, database and networking skills.This will be a varied role involving working on a multitude of cross-platform market-leading technologies to support the running of our bespoke trading platform. The successful candidate will be responsible for all aspects of support covering both proprietary and third-party applications from the front to back office, with a particular focus on Transaction & Regulatory Reporting. Liquidnet champions automation and you will be expected to identify and help streamline manual or repetitive tasks. You will have the opportunity to contribute to, and run with projects, new feature implementations, client migrations, and help Liquidnet migrate to cloud-based technologies. Additionally, the role will involve member user administration and support via phone and email, OMS integration support and trade lifecycle issues for both the MTF platform and trading desk.The successful candidate should possess a positive 'can-do' attitude and an intuitively high level of customer service in their approach. This will complement strong FIX, database (SQL, Sybase or Oracle), as well as Linux, understanding of cloud-based technologies, Windows and Networking troubleshooting skills. Role Responsibilities Contribute towards 'follow the sun' support model, working closely with global teams in APAC and US to ensure pre-market health checks are performed for each region Perform regional start of day health checks to ensure all members are connected to the platform Utilising proprietary tools, provide daily application support and troubleshooting for platform members and internal users, escalating to Development teams appropriately An application support focus on back-office flows, particularly around Regulatory and Transaction Reporting support Daily interaction with all internal stakeholders with regards to support issues Efficiently create and track issues within an incident-management system to help identify trends and patterns Create and monitor internal reports and usage queries Assist with product testing and project work Identify and escalate possible platform improvements Experience / Competences Essential Hands-on support experience within a financial institution (buy-side, sell-side, venue/platform provider) Solid application support experience within a Linux environment Excellent working knowledge of the FIX protocol Good understanding of European Equity market structure, mechanics and flows Ability to convey expected behaviour of industry-standard algorithms (VWAP, TWAP, IS, POV etc) Automation and scripting experience Proven experience of MSSQL, Oracle and Sybase database environments, including complex query-writing Proven experience of supporting Windows Server environments Experience in troubleshooting network problems: i.e. firewall and routing problems Motivated self-starter who takes ownership of responsibilities, and can work autonomously Ability to confidently communicate at all stakeholder levels (technical, client, trader, executive team, etc) Excellent organisational skills Analytical and disciplined approach to problem-solving Must be a team player with ability and interest in participating in new projects and helping other departments within the companyDesired Client / Venue technical FIX onboarding exposure Proven experience in managing cloud-based infrastructure and services, including AWS, Azure, or Google Cloud Platform. Strong understanding of DevOps principles and practices, including CI/CD pipelines, infrastructure as code (IaC), and automated testing Hands-on experience with containerization technologies like Docker and orchestration platforms like Kubernetes. Exposure to supporting message-based architecture Working knowledge of at least one buy-side or sell-side Order Management System Experience with industry-standard monitoring tools (ITRS or similar) Experience with Site Reliability Engineering (SRE) practices, including monitoring, incident response, and post-mortem analysis Job Band & Level Professional / 5 Company Statement We know that the best innovation happens when diverse people with different perspectives and skills work together in an inclusive atmosphere. That's why we're building a culture where everyone plays a part in making people feel welcome, ready and willing to contribute. TP ICAP Accord - our Employee Network - is a central to this. As well as representing specific groups, TP ICAP Accord helps increase awareness, collaboration, shares best practice, and holds our firm to account for driving continuous cultural improvement. Location UK - 135 Bishopsgate - London Connecting clients, communities and colleagues for sustainable growth TP ICAP connects people, platforms, ideas, and insight across the world's financial, energy and commodities markets. As a global leader in market infrastructure and data-led solutions, we enhance market access, increase efficiencies, and unlock possibilities. Work with us Joining TP ICAP puts you at the heart of markets that matter.You'll have the freedom to innovate and act on your initiative. We'll train you and build your abilities in your specialist area, so that you can become an expert in your field. And all within a connected network that's there to set you up for success.TP ICAP Group is a collection of premium brands each with a distinct, client-focused offering. Underpinning and connecting these client-facing brands is the financial security, operational strength and know-how we have as a Group.Connections are at the heart of what we do. We combine our people's know-how with the latest technology to improve price discovery, trade execution and liquidity flow.Connections create strength. Through them, we help our clients to manage risk, realise investment strategies and expand the scope for growth.And connections act as a catalyst. Sparking richer solutions for our clients to break new ground, modernising markets for future performance, and creating dynamic careers for our people. Our capacity to connect builds trust, supports communities and gives us the power to anticipate and respond to change, whatever direction the world takes. It's what makes TP ICAP a mainstay in the global markets, now and in the future.TP ICAP. We connect.
13/07/2026
Full time
Front Office Product Support page is loaded Front Office Product Supportlocations: Londontime type: Full timeposted on: Posted Yesterdayjob requisition id: R5215The TP ICAP Group is a world leading provider of market infrastructure.Our purpose is to provide clients with access to global financial and commodities markets, improving price discovery, liquidity, and distribution of data, through responsible and innovative solutions.Through our people and technology, we connect clients to superior liquidity and data solutions.The Group is home to a stable of premium brands. Collectively, TP ICAP is the largest interdealer broker in the world by revenue, the number one Energy & Commodities broker in the world, the world's leading provider of OTC data, and an award winning all-to-all trading platform.The Group operates from more than 60 offices in 27 countries. We are 5,300 people strong. We work as one to achieve our vision of being the world's most trusted, innovative, liquidity and data solutions specialist. Role Overview Liquidnet is looking for an Application Support engineer to work within the EMEA Front Office Support team. The team requires a motivated self-starter who has the technical skills to support a growing number of buy-side members utilising their FIX, Linux, Windows Server, DevOps, database and networking skills.This will be a varied role involving working on a multitude of cross-platform market-leading technologies to support the running of our bespoke trading platform. The successful candidate will be responsible for all aspects of support covering both proprietary and third-party applications from the front to back office, with a particular focus on Transaction & Regulatory Reporting. Liquidnet champions automation and you will be expected to identify and help streamline manual or repetitive tasks. You will have the opportunity to contribute to, and run with projects, new feature implementations, client migrations, and help Liquidnet migrate to cloud-based technologies. Additionally, the role will involve member user administration and support via phone and email, OMS integration support and trade lifecycle issues for both the MTF platform and trading desk.The successful candidate should possess a positive 'can-do' attitude and an intuitively high level of customer service in their approach. This will complement strong FIX, database (SQL, Sybase or Oracle), as well as Linux, understanding of cloud-based technologies, Windows and Networking troubleshooting skills. Role Responsibilities Contribute towards 'follow the sun' support model, working closely with global teams in APAC and US to ensure pre-market health checks are performed for each region Perform regional start of day health checks to ensure all members are connected to the platform Utilising proprietary tools, provide daily application support and troubleshooting for platform members and internal users, escalating to Development teams appropriately An application support focus on back-office flows, particularly around Regulatory and Transaction Reporting support Daily interaction with all internal stakeholders with regards to support issues Efficiently create and track issues within an incident-management system to help identify trends and patterns Create and monitor internal reports and usage queries Assist with product testing and project work Identify and escalate possible platform improvements Experience / Competences Essential Hands-on support experience within a financial institution (buy-side, sell-side, venue/platform provider) Solid application support experience within a Linux environment Excellent working knowledge of the FIX protocol Good understanding of European Equity market structure, mechanics and flows Ability to convey expected behaviour of industry-standard algorithms (VWAP, TWAP, IS, POV etc) Automation and scripting experience Proven experience of MSSQL, Oracle and Sybase database environments, including complex query-writing Proven experience of supporting Windows Server environments Experience in troubleshooting network problems: i.e. firewall and routing problems Motivated self-starter who takes ownership of responsibilities, and can work autonomously Ability to confidently communicate at all stakeholder levels (technical, client, trader, executive team, etc) Excellent organisational skills Analytical and disciplined approach to problem-solving Must be a team player with ability and interest in participating in new projects and helping other departments within the companyDesired Client / Venue technical FIX onboarding exposure Proven experience in managing cloud-based infrastructure and services, including AWS, Azure, or Google Cloud Platform. Strong understanding of DevOps principles and practices, including CI/CD pipelines, infrastructure as code (IaC), and automated testing Hands-on experience with containerization technologies like Docker and orchestration platforms like Kubernetes. Exposure to supporting message-based architecture Working knowledge of at least one buy-side or sell-side Order Management System Experience with industry-standard monitoring tools (ITRS or similar) Experience with Site Reliability Engineering (SRE) practices, including monitoring, incident response, and post-mortem analysis Job Band & Level Professional / 5 Company Statement We know that the best innovation happens when diverse people with different perspectives and skills work together in an inclusive atmosphere. That's why we're building a culture where everyone plays a part in making people feel welcome, ready and willing to contribute. TP ICAP Accord - our Employee Network - is a central to this. As well as representing specific groups, TP ICAP Accord helps increase awareness, collaboration, shares best practice, and holds our firm to account for driving continuous cultural improvement. Location UK - 135 Bishopsgate - London Connecting clients, communities and colleagues for sustainable growth TP ICAP connects people, platforms, ideas, and insight across the world's financial, energy and commodities markets. As a global leader in market infrastructure and data-led solutions, we enhance market access, increase efficiencies, and unlock possibilities. Work with us Joining TP ICAP puts you at the heart of markets that matter.You'll have the freedom to innovate and act on your initiative. We'll train you and build your abilities in your specialist area, so that you can become an expert in your field. And all within a connected network that's there to set you up for success.TP ICAP Group is a collection of premium brands each with a distinct, client-focused offering. Underpinning and connecting these client-facing brands is the financial security, operational strength and know-how we have as a Group.Connections are at the heart of what we do. We combine our people's know-how with the latest technology to improve price discovery, trade execution and liquidity flow.Connections create strength. Through them, we help our clients to manage risk, realise investment strategies and expand the scope for growth.And connections act as a catalyst. Sparking richer solutions for our clients to break new ground, modernising markets for future performance, and creating dynamic careers for our people. Our capacity to connect builds trust, supports communities and gives us the power to anticipate and respond to change, whatever direction the world takes. It's what makes TP ICAP a mainstay in the global markets, now and in the future.TP ICAP. We connect.
Staff Cloud SRE - AI/ML Platform & GPU Compute
Icehouseventures
The role This is a rare opportunity to be a founding Staff SRE shaping the reliability of large-scale AI systems and GPU compute infrastructure from the ground up. As a Staff Cloud Site Reliability Engineer at Wayve, you will build and scale the reliability foundations of our AI cloud platform. This includes our Model Development Platform (powering end-to-end model development from raw data to on road experimentation) and our GPU Compute platform (large-scale, multi-tenant GPU fleets and scheduling systems driving model training and inference at scale). This is a founding Cloud SRE role. You won't inherit a mature SRE function, you'll help create it. You will define the frameworks, automation, and operational standards that ensure our model development infrastructure, distributed systems, and large compute clusters operate predictably, efficiently, and at scale. This role sits at the intersection of AI research, large-scale cloud infrastructure, and production operations. Your work will directly enable faster model training, reliable experimentation, and scalable AI deployment by ensuring our cloud infrastructure is resilient and performant. Key responsibilities Reliability & Platform Ownership Own the reliability, availability, and performance of the Model Dev Platform and GPU Compute environments. Define and operationalise SLOs, SLIs, and error budgets across platform services. Improve capacity planning, scaling strategies, and resource efficiency across large GPU backed clusters. Partner with ML, platform, and software teams to establish clear production readiness standards. Incident Response & On-Call Participate in a 24/7 on call rotation as first line response for cloud and cluster related incidents. Lead incident triage, escalation, communications, and root cause analysis. Translate post incident learning into durable architectural or automation improvements. Continuously reduce alert noise and recurring operational burden. Observability & Operational Excellence Design and operate monitoring, logging, tracing, and alerting systems that enable rapid detection and recovery. Build dashboards that reflect real user centric platform health (not just infrastructure metrics). Improve deployment safety through better change management, validation, and rollback mechanisms. Automation & Tooling Build automation for cluster operations, training workflows, remediation, and scaling tasks. Implement self healing patterns and resilient recovery workflows. Harden CI/CD and release processes to improve deployment safety and velocity. Support infrastructure as code and policy driven guardrails to ensure secure, reliable cloud environments. About you In order to set you up for success as a Cloud Site Reliability Engineer at Wayve, we're looking for the following skills and experience. Essential skills Proven experience in an SRE, Production Engineer, or Cloud Reliability role supporting large scale cloud systems. Experience operating GPU backed environments or large-scale ML infrastructure. Experience running model training or inference pipelines in production (MLOps). Strong Kubernetes experience, including operating production clusters. Hands on experience running production workloads in AWS, GCP, or Azure. Experience operating complex distributed systems in production, ideally including compute heavy or high performance workloads. Experience working with large compute clusters; exposure to AI/ML training or inference workloads strongly preferred. Strong Linux fundamentals and proficiency in at least one scripting or systems language (e.g. Python, Go, C++) with a bias toward automation. Deep troubleshooting skills across networking, storage, distributed systems, and performance at scale. Experience designing and operating observability stacks (e.g. Datadog, Prometheus, Grafana, OpenTelemetry). Clear communication skills, including leading incidents, writing postmortems, and influencing teams to prioritise reliability improvements. Desirable skills Familiarity with infrastructure as code (e.g. Terraform) and secure cloud production environments. Experience defining and running SLOs/SLIs and building reliability programs across multiple teams. Experience as an early or founding SRE hire establishing processes from scratch. Interest in helping shape and grow a Cloud SRE function, with potential to take on leadership responsibilities over time. Benefits This is a full time role based in our office in London (2 days a week in the office). We operate a hybrid working policy that combines time together in our offices and workshops to fuel innovation, culture, relationships and learning, and time spent working from home. Equal Opportunity Statement Wayve is committed to creating an inclusive interview experience. If you require accommodations or adjustments to participate fully in our interview process, please let us know. We understand that everyone has a unique set of skills and experiences and that not everyone will meet all of the requirements listed above. If you're passionate about self driving cars and think you have what it takes to make a positive impact on the world, we encourage you to apply. At Wayve we're committed to creating a diverse, fair and respectful culture that is inclusive of everyone based on their unique skills and perspectives, and regardless of sex, race, religion or belief, ethnic or national origin, disability, age, citizenship, marital, domestic or civil partnership status, sexual orientation, gender identity, veteran status, pregnancy or related condition (including breastfeeding) or any other basis as protected by applicable law. DISCLAIMER: We will not ask about marriage or pregnancy, care responsibilities or disabilities in any of our job adverts or interviews. However, we do look to capture information about care responsibilities, and disabilities among other diversity information as part of an optional DEI Monitoring form to help us identify areas of improvement in our hiring process and ensure that the process is inclusive and non discriminatory.
13/07/2026
Full time
The role This is a rare opportunity to be a founding Staff SRE shaping the reliability of large-scale AI systems and GPU compute infrastructure from the ground up. As a Staff Cloud Site Reliability Engineer at Wayve, you will build and scale the reliability foundations of our AI cloud platform. This includes our Model Development Platform (powering end-to-end model development from raw data to on road experimentation) and our GPU Compute platform (large-scale, multi-tenant GPU fleets and scheduling systems driving model training and inference at scale). This is a founding Cloud SRE role. You won't inherit a mature SRE function, you'll help create it. You will define the frameworks, automation, and operational standards that ensure our model development infrastructure, distributed systems, and large compute clusters operate predictably, efficiently, and at scale. This role sits at the intersection of AI research, large-scale cloud infrastructure, and production operations. Your work will directly enable faster model training, reliable experimentation, and scalable AI deployment by ensuring our cloud infrastructure is resilient and performant. Key responsibilities Reliability & Platform Ownership Own the reliability, availability, and performance of the Model Dev Platform and GPU Compute environments. Define and operationalise SLOs, SLIs, and error budgets across platform services. Improve capacity planning, scaling strategies, and resource efficiency across large GPU backed clusters. Partner with ML, platform, and software teams to establish clear production readiness standards. Incident Response & On-Call Participate in a 24/7 on call rotation as first line response for cloud and cluster related incidents. Lead incident triage, escalation, communications, and root cause analysis. Translate post incident learning into durable architectural or automation improvements. Continuously reduce alert noise and recurring operational burden. Observability & Operational Excellence Design and operate monitoring, logging, tracing, and alerting systems that enable rapid detection and recovery. Build dashboards that reflect real user centric platform health (not just infrastructure metrics). Improve deployment safety through better change management, validation, and rollback mechanisms. Automation & Tooling Build automation for cluster operations, training workflows, remediation, and scaling tasks. Implement self healing patterns and resilient recovery workflows. Harden CI/CD and release processes to improve deployment safety and velocity. Support infrastructure as code and policy driven guardrails to ensure secure, reliable cloud environments. About you In order to set you up for success as a Cloud Site Reliability Engineer at Wayve, we're looking for the following skills and experience. Essential skills Proven experience in an SRE, Production Engineer, or Cloud Reliability role supporting large scale cloud systems. Experience operating GPU backed environments or large-scale ML infrastructure. Experience running model training or inference pipelines in production (MLOps). Strong Kubernetes experience, including operating production clusters. Hands on experience running production workloads in AWS, GCP, or Azure. Experience operating complex distributed systems in production, ideally including compute heavy or high performance workloads. Experience working with large compute clusters; exposure to AI/ML training or inference workloads strongly preferred. Strong Linux fundamentals and proficiency in at least one scripting or systems language (e.g. Python, Go, C++) with a bias toward automation. Deep troubleshooting skills across networking, storage, distributed systems, and performance at scale. Experience designing and operating observability stacks (e.g. Datadog, Prometheus, Grafana, OpenTelemetry). Clear communication skills, including leading incidents, writing postmortems, and influencing teams to prioritise reliability improvements. Desirable skills Familiarity with infrastructure as code (e.g. Terraform) and secure cloud production environments. Experience defining and running SLOs/SLIs and building reliability programs across multiple teams. Experience as an early or founding SRE hire establishing processes from scratch. Interest in helping shape and grow a Cloud SRE function, with potential to take on leadership responsibilities over time. Benefits This is a full time role based in our office in London (2 days a week in the office). We operate a hybrid working policy that combines time together in our offices and workshops to fuel innovation, culture, relationships and learning, and time spent working from home. Equal Opportunity Statement Wayve is committed to creating an inclusive interview experience. If you require accommodations or adjustments to participate fully in our interview process, please let us know. We understand that everyone has a unique set of skills and experiences and that not everyone will meet all of the requirements listed above. If you're passionate about self driving cars and think you have what it takes to make a positive impact on the world, we encourage you to apply. At Wayve we're committed to creating a diverse, fair and respectful culture that is inclusive of everyone based on their unique skills and perspectives, and regardless of sex, race, religion or belief, ethnic or national origin, disability, age, citizenship, marital, domestic or civil partnership status, sexual orientation, gender identity, veteran status, pregnancy or related condition (including breastfeeding) or any other basis as protected by applicable law. DISCLAIMER: We will not ask about marriage or pregnancy, care responsibilities or disabilities in any of our job adverts or interviews. However, we do look to capture information about care responsibilities, and disabilities among other diversity information as part of an optional DEI Monitoring form to help us identify areas of improvement in our hiring process and ensure that the process is inclusive and non discriminatory.
Senior Cloud & AI Platform Engineer
CreateFuture Leeds, Yorkshire
Position Overview At CreateFuture, the Senior Cloud Engineer is a vital individual contributor and a technical driving force within our delivery teams. You will work closely with clients and delivery teams to bridge the implementation gap, assisting with architectural designs and transforming them into production systems using autonomous and AI-native platforms. This role requires a modern, hybrid engineering approach, blending infrastructure mastery with software development and data engineering principles. You will act as a technical expert on projects, building trusted relationships with client stakeholders, establishing engineering best practices, and mentoring engineers within the Cloud capability. Furthermore, you will actively contribute to setting engineering standards that influence the whole CreateFuture organisation. Key Responsibilities High-Quality Execution: Proactively contribute to project teams by providing technical guidance, resolving complex technical blockages, supporting migration efforts and writing high-quality, spec-driven code. Platform Engineering: Build and maintain independent, self-service developer platforms that allow engineering teams to spin up compliant, ephemer environments automatically using GitOps workflows. AI & Agentic Infrastructure: Implement the data pipelines, workflow orchestrations, and specialised compute footprints needed to support enterprise AI applications, using technologies like AWS Bedrock and AWS AgentCore. Observability & Reliability: Build robust monitoring and observability pipelines to ensure the health, performance, and security of distributed cloud applications and AI models. FinOps Standards: Embed automated cost-estimation tools into the standard CI/CD deployment pipelines, ensuring auto-shutdown policies and resource optimisation are configured by default. Consultancy & Client Alignment Client Advisory: Collaborate with diverse stakeholders, including tech leads, project managers, and enterprise architects, to ensure successful project delivery according to timelines and budgets. Commercial Growth: Have a growth mindset on projects, recognising potential operational bottlenecks or new requirements, and own communicating these throughout the team for new opportunities. Best Practice Champion: Establish best practice tools and processes within client environments, confidently challenging legacy ways of working where appropriate. Capability & Leadership Support Mentoring: Support and mentor engineers within CreateFuture and our client teams, helping them navigate their learning and technical skill gaps. Community Contribution: Actively engage in the internal engineering communities, helping to build out an internal repository of reference architectures, Infrastructure as Code templates, and technical blogs. Culture and Feedback: Help foster an inclusive, collaborative, and psychologically safe team environment by providing timely, open, and honest technical feedback. Skills & Experience Core Technical Capabilities Candidates must demonstrate a hybrid balance across the three strategic pillars: Software Engineering (30%), Infrastructure (40%), and Data Engineering (30%). Infrastructure & Automation (40%): Lots of experience in writing declarative Infrastructure as Code using Terraform, container orchestration with Kubernetes (managing clusters, nodes, and pods), and building CI/CD pipelines via GitHub Actions. Software & Agentic Engineering (30%): Proficiency with modern scripting languages for building custom AI agents, configuring workflow orchestration, connecting enterprise API integrations, and implementing agentic guardrails. Data Engineering (30%): Practical experience implementing components of data pipelines, including real-time streaming tools (AWS Kinesis, Kafka), data orchestration (dbt, Airflow), and managing vector databases for RAG architectures. Observability & Cost Management: SRE/Platform experience with the practical application of real-time monitoring and cloud cost optimisation using native CSP tools or utilities like Infracost and Karpenter. Domain & Sector Experience Regulated Industries: Experience delivering secure platforms within highly regulated environments, such as iGaming, financial services, or banking, where zero trust security and strict compliance controls are mandatory is highly advantageous. Consulting: Proven experience operating in a client-facing or consulting engineering role, demonstrating strong communication skills, relationship-building capabilities, and adaptability across changing technical contexts. Preferred Knowledge-level & Certifications Associate-level cloud certifications (e.g., AWS Solutions Architect Associate or Azure AZ-104). At least one professional-level certification. Progressing toward or holding multiple professional certifications such as CKA (Certified Kubernetes Administrator), AWS ML Engineer Associate, or FinOps Certified Practitioner. What we'll offer you: We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You'll be part of a supportive and friendly culture, surrounded by smart, curious people who care deeply about what they do. We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed. We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.
12/07/2026
Full time
Position Overview At CreateFuture, the Senior Cloud Engineer is a vital individual contributor and a technical driving force within our delivery teams. You will work closely with clients and delivery teams to bridge the implementation gap, assisting with architectural designs and transforming them into production systems using autonomous and AI-native platforms. This role requires a modern, hybrid engineering approach, blending infrastructure mastery with software development and data engineering principles. You will act as a technical expert on projects, building trusted relationships with client stakeholders, establishing engineering best practices, and mentoring engineers within the Cloud capability. Furthermore, you will actively contribute to setting engineering standards that influence the whole CreateFuture organisation. Key Responsibilities High-Quality Execution: Proactively contribute to project teams by providing technical guidance, resolving complex technical blockages, supporting migration efforts and writing high-quality, spec-driven code. Platform Engineering: Build and maintain independent, self-service developer platforms that allow engineering teams to spin up compliant, ephemer environments automatically using GitOps workflows. AI & Agentic Infrastructure: Implement the data pipelines, workflow orchestrations, and specialised compute footprints needed to support enterprise AI applications, using technologies like AWS Bedrock and AWS AgentCore. Observability & Reliability: Build robust monitoring and observability pipelines to ensure the health, performance, and security of distributed cloud applications and AI models. FinOps Standards: Embed automated cost-estimation tools into the standard CI/CD deployment pipelines, ensuring auto-shutdown policies and resource optimisation are configured by default. Consultancy & Client Alignment Client Advisory: Collaborate with diverse stakeholders, including tech leads, project managers, and enterprise architects, to ensure successful project delivery according to timelines and budgets. Commercial Growth: Have a growth mindset on projects, recognising potential operational bottlenecks or new requirements, and own communicating these throughout the team for new opportunities. Best Practice Champion: Establish best practice tools and processes within client environments, confidently challenging legacy ways of working where appropriate. Capability & Leadership Support Mentoring: Support and mentor engineers within CreateFuture and our client teams, helping them navigate their learning and technical skill gaps. Community Contribution: Actively engage in the internal engineering communities, helping to build out an internal repository of reference architectures, Infrastructure as Code templates, and technical blogs. Culture and Feedback: Help foster an inclusive, collaborative, and psychologically safe team environment by providing timely, open, and honest technical feedback. Skills & Experience Core Technical Capabilities Candidates must demonstrate a hybrid balance across the three strategic pillars: Software Engineering (30%), Infrastructure (40%), and Data Engineering (30%). Infrastructure & Automation (40%): Lots of experience in writing declarative Infrastructure as Code using Terraform, container orchestration with Kubernetes (managing clusters, nodes, and pods), and building CI/CD pipelines via GitHub Actions. Software & Agentic Engineering (30%): Proficiency with modern scripting languages for building custom AI agents, configuring workflow orchestration, connecting enterprise API integrations, and implementing agentic guardrails. Data Engineering (30%): Practical experience implementing components of data pipelines, including real-time streaming tools (AWS Kinesis, Kafka), data orchestration (dbt, Airflow), and managing vector databases for RAG architectures. Observability & Cost Management: SRE/Platform experience with the practical application of real-time monitoring and cloud cost optimisation using native CSP tools or utilities like Infracost and Karpenter. Domain & Sector Experience Regulated Industries: Experience delivering secure platforms within highly regulated environments, such as iGaming, financial services, or banking, where zero trust security and strict compliance controls are mandatory is highly advantageous. Consulting: Proven experience operating in a client-facing or consulting engineering role, demonstrating strong communication skills, relationship-building capabilities, and adaptability across changing technical contexts. Preferred Knowledge-level & Certifications Associate-level cloud certifications (e.g., AWS Solutions Architect Associate or Azure AZ-104). At least one professional-level certification. Progressing toward or holding multiple professional certifications such as CKA (Certified Kubernetes Administrator), AWS ML Engineer Associate, or FinOps Certified Practitioner. What we'll offer you: We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You'll be part of a supportive and friendly culture, surrounded by smart, curious people who care deeply about what they do. We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed. We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.
Senior Cloud & AI Platform Engineer
CreateFuture
Position Overview At CreateFuture, the Senior Cloud Engineer is a vital individual contributor and a technical driving force within our delivery teams. You will work closely with clients and delivery teams to bridge the implementation gap, assisting with architectural designs and transforming them into production systems using autonomous and AI-native platforms. This role requires a modern, hybrid engineering approach, blending infrastructure mastery with software development and data engineering principles. You will act as a technical expert on projects, building trusted relationships with client stakeholders, establishing engineering best practices, and mentoring engineers within the Cloud capability. Furthermore, you will actively contribute to setting engineering standards that influence the whole CreateFuture organisation. Key Responsibilities High-Quality Execution: Proactively contribute to project teams by providing technical guidance, resolving complex technical blockages, supporting migration efforts and writing high-quality, spec-driven code. Platform Engineering: Build and maintain independent, self-service developer platforms that allow engineering teams to spin up compliant, ephemer environments automatically using GitOps workflows. AI & Agentic Infrastructure: Implement the data pipelines, workflow orchestrations, and specialised compute footprints needed to support enterprise AI applications, using technologies like AWS Bedrock and AWS AgentCore. Observability & Reliability: Build robust monitoring and observability pipelines to ensure the health, performance, and security of distributed cloud applications and AI models. FinOps Standards: Embed automated cost-estimation tools into the standard CI/CD deployment pipelines, ensuring auto-shutdown policies and resource optimisation are configured by default. Consultancy & Client Alignment Client Advisory: Collaborate with diverse stakeholders, including tech leads, project managers, and enterprise architects, to ensure successful project delivery according to timelines and budgets. Commercial Growth: Have a growth mindset on projects, recognising potential operational bottlenecks or new requirements, and own communicating these throughout the team for new opportunities. Best Practice Champion: Establish best practice tools and processes within client environments, confidently challenging legacy ways of working where appropriate. Capability & Leadership Support Mentoring: Support and mentor engineers within CreateFuture and our client teams, helping them navigate their learning and technical skill gaps. Community Contribution: Actively engage in the internal engineering communities, helping to build out an internal repository of reference architectures, Infrastructure as Code templates, and technical blogs. Culture and Feedback: Help foster an inclusive, collaborative, and psychologically safe team environment by providing timely, open, and honest technical feedback. Skills & Experience Core Technical Capabilities Candidates must demonstrate a hybrid balance across the three strategic pillars: Software Engineering (30%), Infrastructure (40%), and Data Engineering (30%). Infrastructure & Automation (40%): Lots of experience in writing declarative Infrastructure as Code using Terraform, container orchestration with Kubernetes (managing clusters, nodes, and pods), and building CI/CD pipelines via GitHub Actions. Software & Agentic Engineering (30%): Proficiency with modern scripting languages for building custom AI agents, configuring workflow orchestration, connecting enterprise API integrations, and implementing agentic guardrails. Data Engineering (30%): Practical experience implementing components of data pipelines, including real-time streaming tools (AWS Kinesis, Kafka), data orchestration (dbt, Airflow), and managing vector databases for RAG architectures. Observability & Cost Management: SRE/Platform experience with the practical application of real-time monitoring and cloud cost optimisation using native CSP tools or utilities like Infracost and Karpenter. Domain & Sector Experience Regulated Industries: Experience delivering secure platforms within highly regulated environments, such as iGaming, financial services, or banking, where zero trust security and strict compliance controls are mandatory is highly advantageous. Consulting: Proven experience operating in a client-facing or consulting engineering role, demonstrating strong communication skills, relationship-building capabilities, and adaptability across changing technical contexts. Preferred Knowledge-level & Certifications Associate-level cloud certifications (e.g., AWS Solutions Architect Associate or Azure AZ-104). At least one professional-level certification. Progressing toward or holding multiple professional certifications such as CKA (Certified Kubernetes Administrator), AWS ML Engineer Associate, or FinOps Certified Practitioner. What we'll offer you: We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You'll be part of a supportive and friendly culture, surrounded by smart, curious people who care deeply about what they do. We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed. We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.
12/07/2026
Full time
Position Overview At CreateFuture, the Senior Cloud Engineer is a vital individual contributor and a technical driving force within our delivery teams. You will work closely with clients and delivery teams to bridge the implementation gap, assisting with architectural designs and transforming them into production systems using autonomous and AI-native platforms. This role requires a modern, hybrid engineering approach, blending infrastructure mastery with software development and data engineering principles. You will act as a technical expert on projects, building trusted relationships with client stakeholders, establishing engineering best practices, and mentoring engineers within the Cloud capability. Furthermore, you will actively contribute to setting engineering standards that influence the whole CreateFuture organisation. Key Responsibilities High-Quality Execution: Proactively contribute to project teams by providing technical guidance, resolving complex technical blockages, supporting migration efforts and writing high-quality, spec-driven code. Platform Engineering: Build and maintain independent, self-service developer platforms that allow engineering teams to spin up compliant, ephemer environments automatically using GitOps workflows. AI & Agentic Infrastructure: Implement the data pipelines, workflow orchestrations, and specialised compute footprints needed to support enterprise AI applications, using technologies like AWS Bedrock and AWS AgentCore. Observability & Reliability: Build robust monitoring and observability pipelines to ensure the health, performance, and security of distributed cloud applications and AI models. FinOps Standards: Embed automated cost-estimation tools into the standard CI/CD deployment pipelines, ensuring auto-shutdown policies and resource optimisation are configured by default. Consultancy & Client Alignment Client Advisory: Collaborate with diverse stakeholders, including tech leads, project managers, and enterprise architects, to ensure successful project delivery according to timelines and budgets. Commercial Growth: Have a growth mindset on projects, recognising potential operational bottlenecks or new requirements, and own communicating these throughout the team for new opportunities. Best Practice Champion: Establish best practice tools and processes within client environments, confidently challenging legacy ways of working where appropriate. Capability & Leadership Support Mentoring: Support and mentor engineers within CreateFuture and our client teams, helping them navigate their learning and technical skill gaps. Community Contribution: Actively engage in the internal engineering communities, helping to build out an internal repository of reference architectures, Infrastructure as Code templates, and technical blogs. Culture and Feedback: Help foster an inclusive, collaborative, and psychologically safe team environment by providing timely, open, and honest technical feedback. Skills & Experience Core Technical Capabilities Candidates must demonstrate a hybrid balance across the three strategic pillars: Software Engineering (30%), Infrastructure (40%), and Data Engineering (30%). Infrastructure & Automation (40%): Lots of experience in writing declarative Infrastructure as Code using Terraform, container orchestration with Kubernetes (managing clusters, nodes, and pods), and building CI/CD pipelines via GitHub Actions. Software & Agentic Engineering (30%): Proficiency with modern scripting languages for building custom AI agents, configuring workflow orchestration, connecting enterprise API integrations, and implementing agentic guardrails. Data Engineering (30%): Practical experience implementing components of data pipelines, including real-time streaming tools (AWS Kinesis, Kafka), data orchestration (dbt, Airflow), and managing vector databases for RAG architectures. Observability & Cost Management: SRE/Platform experience with the practical application of real-time monitoring and cloud cost optimisation using native CSP tools or utilities like Infracost and Karpenter. Domain & Sector Experience Regulated Industries: Experience delivering secure platforms within highly regulated environments, such as iGaming, financial services, or banking, where zero trust security and strict compliance controls are mandatory is highly advantageous. Consulting: Proven experience operating in a client-facing or consulting engineering role, demonstrating strong communication skills, relationship-building capabilities, and adaptability across changing technical contexts. Preferred Knowledge-level & Certifications Associate-level cloud certifications (e.g., AWS Solutions Architect Associate or Azure AZ-104). At least one professional-level certification. Progressing toward or holding multiple professional certifications such as CKA (Certified Kubernetes Administrator), AWS ML Engineer Associate, or FinOps Certified Practitioner. What we'll offer you: We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You'll be part of a supportive and friendly culture, surrounded by smart, curious people who care deeply about what they do. We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed. We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.
Lead SRE - Data Product Studio
JPMorgan Chase & Co.
Build systems that stay available when it matters most. In this role, you'll shape how reliability is engineered across critical data products and platforms at JPMorganChase. You'll partner with engineers and stakeholders to set measurable service goals, reduce toil through automation, and raise operational standards. You'll have room to lead, mentor, and influence technical direction while tackling complex, high-impact challenges. If you're passionate about scalable, secure, and resilient systems, you'll thrive here. As a Lead Site Reliability Engineer in Corporate Technology: Data Strategy & Architecture you will lead resiliency design reviews and champion site reliability practices across medium to large-sized products. You will break down complex problems into actionable work, guide delivery through strong engineering standards, and help teams improve reliability using data-driven insights. You will serve as a technical leader during major incidents and influence solutions across multiple technical domains. You will mentor engineers and advise on technical and business issues to improve outcomes for customers and stakeholders. This is a technical leadership (IC) position, rather than a team leadership position. Job Responsibilities Champion site reliability culture and practices, exerting technical influence across the team Lead initiatives to improve reliability and stability using data-driven analytics to improve service levels Partner with engineers and stakeholders to define service level indicators, service level objectives, and error budgets Identify and resolve technical bottlenecks across one or more technical domains Act as a primary point of contact during major incidents, driving rapid diagnosis and resolution to reduce impact Document and share knowledge through internal forums and communities of practice Define and drive CI/CD strategy and pipeline standards, including artifact management, environment promotion, and release gating Design and implement infrastructure-as-code practices to enable self-service provisioning and reduce toil Lead resiliency design reviews for new and existing data products, identifying single points of failure, capacity risks, and operational gaps before production Embed shift-left security practices into delivery pipelines, including automated scanning, secrets management, and policy-as-code Lead capacity planning and cost optimization across cloud and on-premise compute for data workloads Required Qualifications, Capabilities, and Skills Advanced knowledge of site reliability engineering practices, including reliability, scalability, performance, security, enterprise architecture, and toil reduction Proficiency in at least one programming language (e.g., Python, Java Spring Boot, .NET) Experience designing, implementing, and operating observability practices (e.g., white/black box monitoring, SLO alerting, telemetry collection) using tools such as Grafana, Dynatrace, Prometheus, Datadog, or Splunk Hands-on experience designing and operating CI/CD pipelines and release automation (e.g., Jenkins, GitLab CI, GitHub Actions, Argo CD) Proficiency in infrastructure-as-code and configuration management tooling (e.g., Terraform, Ansible, Helm) Experience with containers and orchestration (e.g., Docker, Kubernetes, ECS) Ability to troubleshoot common networking technologies and issues Ability to identify and solve problems involving complex data structures and algorithms Ability to break down complex problems into clear, deliverable work for engineering teams Ability to collaborate across stakeholder groups and influence technical decisions Drive to self-educate and evaluate new technology Preferred Qualifications, Capabilities, and Skills Experience in a data platform, data mesh, or data product engineering environment (e.g., Spark, Kafka, dbt, Airflow) Familiarity with cloud-native data services on AWS, Azure, or GCP Experience with GitOps workflows and progressive delivery patterns Knowledge of chaos engineering practices and tooling Experience mentoring engineers and driving engineering standards across teams Relevant certifications (e.g., CKA, AWS/GCP/Azure Solutions Architect, HashiCorp Terraform Associate)
12/07/2026
Full time
Build systems that stay available when it matters most. In this role, you'll shape how reliability is engineered across critical data products and platforms at JPMorganChase. You'll partner with engineers and stakeholders to set measurable service goals, reduce toil through automation, and raise operational standards. You'll have room to lead, mentor, and influence technical direction while tackling complex, high-impact challenges. If you're passionate about scalable, secure, and resilient systems, you'll thrive here. As a Lead Site Reliability Engineer in Corporate Technology: Data Strategy & Architecture you will lead resiliency design reviews and champion site reliability practices across medium to large-sized products. You will break down complex problems into actionable work, guide delivery through strong engineering standards, and help teams improve reliability using data-driven insights. You will serve as a technical leader during major incidents and influence solutions across multiple technical domains. You will mentor engineers and advise on technical and business issues to improve outcomes for customers and stakeholders. This is a technical leadership (IC) position, rather than a team leadership position. Job Responsibilities Champion site reliability culture and practices, exerting technical influence across the team Lead initiatives to improve reliability and stability using data-driven analytics to improve service levels Partner with engineers and stakeholders to define service level indicators, service level objectives, and error budgets Identify and resolve technical bottlenecks across one or more technical domains Act as a primary point of contact during major incidents, driving rapid diagnosis and resolution to reduce impact Document and share knowledge through internal forums and communities of practice Define and drive CI/CD strategy and pipeline standards, including artifact management, environment promotion, and release gating Design and implement infrastructure-as-code practices to enable self-service provisioning and reduce toil Lead resiliency design reviews for new and existing data products, identifying single points of failure, capacity risks, and operational gaps before production Embed shift-left security practices into delivery pipelines, including automated scanning, secrets management, and policy-as-code Lead capacity planning and cost optimization across cloud and on-premise compute for data workloads Required Qualifications, Capabilities, and Skills Advanced knowledge of site reliability engineering practices, including reliability, scalability, performance, security, enterprise architecture, and toil reduction Proficiency in at least one programming language (e.g., Python, Java Spring Boot, .NET) Experience designing, implementing, and operating observability practices (e.g., white/black box monitoring, SLO alerting, telemetry collection) using tools such as Grafana, Dynatrace, Prometheus, Datadog, or Splunk Hands-on experience designing and operating CI/CD pipelines and release automation (e.g., Jenkins, GitLab CI, GitHub Actions, Argo CD) Proficiency in infrastructure-as-code and configuration management tooling (e.g., Terraform, Ansible, Helm) Experience with containers and orchestration (e.g., Docker, Kubernetes, ECS) Ability to troubleshoot common networking technologies and issues Ability to identify and solve problems involving complex data structures and algorithms Ability to break down complex problems into clear, deliverable work for engineering teams Ability to collaborate across stakeholder groups and influence technical decisions Drive to self-educate and evaluate new technology Preferred Qualifications, Capabilities, and Skills Experience in a data platform, data mesh, or data product engineering environment (e.g., Spark, Kafka, dbt, Airflow) Familiarity with cloud-native data services on AWS, Azure, or GCP Experience with GitOps workflows and progressive delivery patterns Knowledge of chaos engineering practices and tooling Experience mentoring engineers and driving engineering standards across teams Relevant certifications (e.g., CKA, AWS/GCP/Azure Solutions Architect, HashiCorp Terraform Associate)
Senior Cloud & AI Platform Engineer
CreateFuture Edinburgh, Midlothian
Position Overview At CreateFuture, the Senior Cloud Engineer is a vital individual contributor and a technical driving force within our delivery teams. You will work closely with clients and delivery teams to bridge the implementation gap, assisting with architectural designs and transforming them into production systems using autonomous and AI-native platforms. This role requires a modern, hybrid engineering approach, blending infrastructure mastery with software development and data engineering principles. You will act as a technical expert on projects, building trusted relationships with client stakeholders, establishing engineering best practices, and mentoring engineers within the Cloud capability. Furthermore, you will actively contribute to setting engineering standards that influence the whole CreateFuture organisation. Key Responsibilities High-Quality Execution: Proactively contribute to project teams by providing technical guidance, resolving complex technical blockages, supporting migration efforts and writing high-quality, spec-driven code. Platform Engineering: Build and maintain independent, self-service developer platforms that allow engineering teams to spin up compliant, ephemer environments automatically using GitOps workflows. AI & Agentic Infrastructure: Implement the data pipelines, workflow orchestrations, and specialised compute footprints needed to support enterprise AI applications, using technologies like AWS Bedrock and AWS AgentCore. Observability & Reliability: Build robust monitoring and observability pipelines to ensure the health, performance, and security of distributed cloud applications and AI models. FinOps Standards: Embed automated cost-estimation tools into the standard CI/CD deployment pipelines, ensuring auto-shutdown policies and resource optimisation are configured by default. Consultancy & Client Alignment Client Advisory: Collaborate with diverse stakeholders, including tech leads, project managers, and enterprise architects, to ensure successful project delivery according to timelines and budgets. Commercial Growth: Have a growth mindset on projects, recognising potential operational bottlenecks or new requirements, and own communicating these throughout the team for new opportunities. Best Practice Champion: Establish best practice tools and processes within client environments, confidently challenging legacy ways of working where appropriate. Capability & Leadership Support Mentoring: Support and mentor engineers within CreateFuture and our client teams, helping them navigate their learning and technical skill gaps. Community Contribution: Actively engage in the internal engineering communities, helping to build out an internal repository of reference architectures, Infrastructure as Code templates, and technical blogs. Culture and Feedback: Help foster an inclusive, collaborative, and psychologically safe team environment by providing timely, open, and honest technical feedback. Skills & Experience Core Technical Capabilities Candidates must demonstrate a hybrid balance across the three strategic pillars: Software Engineering (30%), Infrastructure (40%), and Data Engineering (30%). Infrastructure & Automation (40%): Lots of experience in writing declarative Infrastructure as Code using Terraform, container orchestration with Kubernetes (managing clusters, nodes, and pods), and building CI/CD pipelines via GitHub Actions. Software & Agentic Engineering (30%): Proficiency with modern scripting languages for building custom AI agents, configuring workflow orchestration, connecting enterprise API integrations, and implementing agentic guardrails. Data Engineering (30%): Practical experience implementing components of data pipelines, including real-time streaming tools (AWS Kinesis, Kafka), data orchestration (dbt, Airflow), and managing vector databases for RAG architectures. Observability & Cost Management: SRE/Platform experience with the practical application of real-time monitoring and cloud cost optimisation using native CSP tools or utilities like Infracost and Karpenter. Domain & Sector Experience Regulated Industries: Experience delivering secure platforms within highly regulated environments, such as iGaming, financial services, or banking, where zero trust security and strict compliance controls are mandatory is highly advantageous. Consulting: Proven experience operating in a client-facing or consulting engineering role, demonstrating strong communication skills, relationship-building capabilities, and adaptability across changing technical contexts. Preferred Knowledge-level & Certifications Associate-level cloud certifications (e.g., AWS Solutions Architect Associate or Azure AZ-104). At least one professional-level certification. Progressing toward or holding multiple professional certifications such as CKA (Certified Kubernetes Administrator), AWS ML Engineer Associate, or FinOps Certified Practitioner. What we'll offer you: We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You'll be part of a supportive and friendly culture, surrounded by smart, curious people who care deeply about what they do. We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed. We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.
12/07/2026
Full time
Position Overview At CreateFuture, the Senior Cloud Engineer is a vital individual contributor and a technical driving force within our delivery teams. You will work closely with clients and delivery teams to bridge the implementation gap, assisting with architectural designs and transforming them into production systems using autonomous and AI-native platforms. This role requires a modern, hybrid engineering approach, blending infrastructure mastery with software development and data engineering principles. You will act as a technical expert on projects, building trusted relationships with client stakeholders, establishing engineering best practices, and mentoring engineers within the Cloud capability. Furthermore, you will actively contribute to setting engineering standards that influence the whole CreateFuture organisation. Key Responsibilities High-Quality Execution: Proactively contribute to project teams by providing technical guidance, resolving complex technical blockages, supporting migration efforts and writing high-quality, spec-driven code. Platform Engineering: Build and maintain independent, self-service developer platforms that allow engineering teams to spin up compliant, ephemer environments automatically using GitOps workflows. AI & Agentic Infrastructure: Implement the data pipelines, workflow orchestrations, and specialised compute footprints needed to support enterprise AI applications, using technologies like AWS Bedrock and AWS AgentCore. Observability & Reliability: Build robust monitoring and observability pipelines to ensure the health, performance, and security of distributed cloud applications and AI models. FinOps Standards: Embed automated cost-estimation tools into the standard CI/CD deployment pipelines, ensuring auto-shutdown policies and resource optimisation are configured by default. Consultancy & Client Alignment Client Advisory: Collaborate with diverse stakeholders, including tech leads, project managers, and enterprise architects, to ensure successful project delivery according to timelines and budgets. Commercial Growth: Have a growth mindset on projects, recognising potential operational bottlenecks or new requirements, and own communicating these throughout the team for new opportunities. Best Practice Champion: Establish best practice tools and processes within client environments, confidently challenging legacy ways of working where appropriate. Capability & Leadership Support Mentoring: Support and mentor engineers within CreateFuture and our client teams, helping them navigate their learning and technical skill gaps. Community Contribution: Actively engage in the internal engineering communities, helping to build out an internal repository of reference architectures, Infrastructure as Code templates, and technical blogs. Culture and Feedback: Help foster an inclusive, collaborative, and psychologically safe team environment by providing timely, open, and honest technical feedback. Skills & Experience Core Technical Capabilities Candidates must demonstrate a hybrid balance across the three strategic pillars: Software Engineering (30%), Infrastructure (40%), and Data Engineering (30%). Infrastructure & Automation (40%): Lots of experience in writing declarative Infrastructure as Code using Terraform, container orchestration with Kubernetes (managing clusters, nodes, and pods), and building CI/CD pipelines via GitHub Actions. Software & Agentic Engineering (30%): Proficiency with modern scripting languages for building custom AI agents, configuring workflow orchestration, connecting enterprise API integrations, and implementing agentic guardrails. Data Engineering (30%): Practical experience implementing components of data pipelines, including real-time streaming tools (AWS Kinesis, Kafka), data orchestration (dbt, Airflow), and managing vector databases for RAG architectures. Observability & Cost Management: SRE/Platform experience with the practical application of real-time monitoring and cloud cost optimisation using native CSP tools or utilities like Infracost and Karpenter. Domain & Sector Experience Regulated Industries: Experience delivering secure platforms within highly regulated environments, such as iGaming, financial services, or banking, where zero trust security and strict compliance controls are mandatory is highly advantageous. Consulting: Proven experience operating in a client-facing or consulting engineering role, demonstrating strong communication skills, relationship-building capabilities, and adaptability across changing technical contexts. Preferred Knowledge-level & Certifications Associate-level cloud certifications (e.g., AWS Solutions Architect Associate or Azure AZ-104). At least one professional-level certification. Progressing toward or holding multiple professional certifications such as CKA (Certified Kubernetes Administrator), AWS ML Engineer Associate, or FinOps Certified Practitioner. What we'll offer you: We trust people to do their best work. That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally. You'll be part of a supportive and friendly culture, surrounded by smart, curious people who care deeply about what they do. We offer flexible working, including hybrid and remote options. Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or CreateFuture offices when needed. We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.
Senior Site Reliability Engineer
P2P
Senior Site Reliability Engineer (SRE) - GCP/Kubernetes We are seeking an experienced and highly motivated Senior Site Reliability Engineer (SRE) to join our small, agile engineering team. This role offers the unique opportunity to drive the reliability, scalability, and performance of our core platform with a high degree of autonomy and ownership. The successful candidate will split their time between providing expert operational support for our critical systems and leading exciting new infrastructure projects. Our mindset is to get the right person not the person with the skills that match our stack. It is important to be able to foresee problems before they show up and create solutions that mitigate them. If you enjoy a challenging environment, implementing "infrastructure as code" principles, and directly seeing the impact of your work, this is the place for you. What You Will Do Design & Build: Architect, deploy, and maintain highly scalable and reliable infrastructure on Google Cloud Platform (GCP) using Kubernetes and Infrastructure-as-Code tools. Automation: Champion automation across the entire software development lifecycle (SDLC), utilizing IaC, Python and Bash to reduce toil and improve operational efficiency. Infrastructure-as-Code (IaC): Own and evolve our declarative infrastructure using Terraform for cloud resources and Helm for Kubernetes application deployment. Monitoring & Observability: Implement and manage robust monitoring, alerting, and logging solutions to ensure clear system visibility and proactive issue identification. Reliability & Performance: Define, measure, and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Participate in on-call rotation (if applicable) and lead post-incident reviews to drive continuous improvement. Collaboration: Work closely with software development teams to provide expert guidance on deployment strategies, scalability concerns, and cloud-native best practices. Ownership: Take full ownership of projects from inception through to production operation, including documentation and knowledge transfer. Required Experience & Skills Core Technical Stack Cloud Platform: 4+ years of hands on experience with Google Cloud Platform (GCP) (or similar cloud infrastructure). Container Orchestration: Expert level proficiency in managing, scaling, and troubleshooting production Kubernetes environments. Infrastructure-as-Code: Deep expertise in Terraform for managing cloud and Kubernetes resources. Deployment: Strong experience with Helm for packaging and deploying applications on Kubernetes. Scripting/Programming: Proficient in at least one major programming language, preferably Python, for automation and tool development. Tooling & Concepts CI/CD: Experience setting up and maintaining modern CI/CD pipelines. Observability: Practical experience implementing and managing monitoring and logging tools. Networking: Solid understanding of TCP/IP, load balancing, DNS, and cloud native networking within Kubernetes. Operating Systems: Strong command line skills and experience with Linux systems. Soft Skills & Team Fit High Ownership: Demonstrated ability to own a problem end to end, from investigation to resolution and preventative measures. Small Team Mentality: Happy to be a generalist and switch context quickly between support tickets, operational toil reduction, and long term project work. Adaptability: A proven track record of rapidly learning and applying new technologies and tools. Equivalent experience with other clouds (AWS/Azure) or similar tools is highly valued. Communication: Excellent verbal and written communication skills for documentation and interacting with non technical stakeholders. Bonus Points For Familiarity with Service Mesh technologies (e.g., Istio). Experience in security best practices within cloud and container environments (e.g., hardening, secrets management). Certifications in GCP or Kubernetes (e.g., CKAD, CKA, Professional Cloud DevOps Engineer).
12/07/2026
Full time
Senior Site Reliability Engineer (SRE) - GCP/Kubernetes We are seeking an experienced and highly motivated Senior Site Reliability Engineer (SRE) to join our small, agile engineering team. This role offers the unique opportunity to drive the reliability, scalability, and performance of our core platform with a high degree of autonomy and ownership. The successful candidate will split their time between providing expert operational support for our critical systems and leading exciting new infrastructure projects. Our mindset is to get the right person not the person with the skills that match our stack. It is important to be able to foresee problems before they show up and create solutions that mitigate them. If you enjoy a challenging environment, implementing "infrastructure as code" principles, and directly seeing the impact of your work, this is the place for you. What You Will Do Design & Build: Architect, deploy, and maintain highly scalable and reliable infrastructure on Google Cloud Platform (GCP) using Kubernetes and Infrastructure-as-Code tools. Automation: Champion automation across the entire software development lifecycle (SDLC), utilizing IaC, Python and Bash to reduce toil and improve operational efficiency. Infrastructure-as-Code (IaC): Own and evolve our declarative infrastructure using Terraform for cloud resources and Helm for Kubernetes application deployment. Monitoring & Observability: Implement and manage robust monitoring, alerting, and logging solutions to ensure clear system visibility and proactive issue identification. Reliability & Performance: Define, measure, and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Participate in on-call rotation (if applicable) and lead post-incident reviews to drive continuous improvement. Collaboration: Work closely with software development teams to provide expert guidance on deployment strategies, scalability concerns, and cloud-native best practices. Ownership: Take full ownership of projects from inception through to production operation, including documentation and knowledge transfer. Required Experience & Skills Core Technical Stack Cloud Platform: 4+ years of hands on experience with Google Cloud Platform (GCP) (or similar cloud infrastructure). Container Orchestration: Expert level proficiency in managing, scaling, and troubleshooting production Kubernetes environments. Infrastructure-as-Code: Deep expertise in Terraform for managing cloud and Kubernetes resources. Deployment: Strong experience with Helm for packaging and deploying applications on Kubernetes. Scripting/Programming: Proficient in at least one major programming language, preferably Python, for automation and tool development. Tooling & Concepts CI/CD: Experience setting up and maintaining modern CI/CD pipelines. Observability: Practical experience implementing and managing monitoring and logging tools. Networking: Solid understanding of TCP/IP, load balancing, DNS, and cloud native networking within Kubernetes. Operating Systems: Strong command line skills and experience with Linux systems. Soft Skills & Team Fit High Ownership: Demonstrated ability to own a problem end to end, from investigation to resolution and preventative measures. Small Team Mentality: Happy to be a generalist and switch context quickly between support tickets, operational toil reduction, and long term project work. Adaptability: A proven track record of rapidly learning and applying new technologies and tools. Equivalent experience with other clouds (AWS/Azure) or similar tools is highly valued. Communication: Excellent verbal and written communication skills for documentation and interacting with non technical stakeholders. Bonus Points For Familiarity with Service Mesh technologies (e.g., Istio). Experience in security best practices within cloud and container environments (e.g., hardening, secrets management). Certifications in GCP or Kubernetes (e.g., CKAD, CKA, Professional Cloud DevOps Engineer).
London Stock Exchange Group
Software & Data Engineers
London Stock Exchange Group
Software & Data EngineersApplylocations: London, United Kingdomtime type: Full timeposted on: Posted Todayjob requisition id: R ROLE SUMMARY: FTSE Russell, part of LSEG, powers some of the world's most recognised benchmarks and index solutions. Our technology, data and products help investors measure, manage and capture market opportunities across global financial markets. As we continue to modernise our platforms and scale our capabilities, our engineers, product specialists and data experts play a critical role in delivering innovative solutions that support investment decisions for clients worldwide. WHAT YOU'LL DO: We're looking for a hands-on engineers to help modernise and scale the FTSE platform - a cloud-native, low-latency system that operates at global scale and underpins critical investment products used worldwide.We are recruiting for individual contributors across a varied technical tech. You'll work on challenging problems across software engineering where performance, data quality, and reliability are non-negotiable, while leveraging modern cloud (AWS) and AI-assisted development tooling to accelerate delivery without compromising control.You will be need to have practical use of AI-assisted development tooling to improve engineering productivity. WHAT YOU'LL BRING: Experience developing software using Java, Python, C++, C#, or similar technologies Experience with cloud platforms such as AWS or Azure Knowledge of CI/CD, automation, testing, and modern engineering practices Strong problem-solving and analytical skills Interest in AI and emerging technologies Ability to collaborate effectively across global teams Bachelor's degree in computer science, Engineering, or equivalent practical experience Experience We Value We value curiosity, innovation, and continuous learning, and we're looking for engineers who are excited by emerging technologies and motivated to make a meaningful impact . Join us and help build the technology and data platforms that sit at the heart of global finance. We welcome professionals with experience in one or more of the following areas: Software Engineering Data Engineering Cloud & Platform Engineering Site Reliability Engineering (SRE) AI & Machine Learning DevOps & Automation Architecture & Distributed Systems Analytics & Data PlatformsBring your curiosity, expertise, and ambition-and help build what's next for global investment and market intelligence. ABOUT US: LSEG (London Stock Exchange Group) is more than a diversified global financial markets infrastructure and data business. We are dedicated, open-access partners with a dedication to excellence in delivering the services our customers expect from us. With extensive experience, deep knowledge and worldwide presence across financial markets, we enable businesses and economies around the world to fund innovation, manage risk and create jobs. It's how we've contributed to supporting the financial stability and growth of communities and economies globally for more than 300 years. Through a comprehensive suite of trusted financial market infrastructure services - and our open-access model - we provide the flexibility, stability and trust that enable our customers to pursue their ambitions with confidence and clarity.LSEG is headquartered in the United Kingdom, with significant operations in 70 countries across EMEA, North America, Latin America and Asia Pacific. We employ 25,000 people globally, more than half located in Asia Pacific. LSEG's ticker symbol is LSEG. OUR PEOPLE: People are at the heart of what we do and drive the success of our business. Our culture of connecting, creating opportunity and delivering excellence shape how we think, how we do things and how we help our people fulfil their potential. We embrace diversity and actively seek to attract individuals with unique backgrounds and perspectives. We break down barriers and encourage teamwork, enabling innovation and rapid development of solutions that make a difference. Our workplace generates an enriching and rewarding experience for our people and customers alike. Our vision is to build an inclusive culture in which everyone feels encouraged to fulfil their potential.We know that real personal growth cannot be achieved by simply climbing a career ladder - which is why we encourage and enable a wealth of avenues and interesting opportunities for everyone to broaden and deepen their skills and expertise. As a global organisation spanning 70 countries and one rooted in a culture of growth, opportunity, diversity and innovation, LSEG is a place where everyone can grow, develop and fulfil your potential with meaningful careers Career Stage: Senior Associate London Stock Exchange Group (LSEG) Information: Join us and be part of a team that values innovation, quality, and continuous improvement. If you're ready to take your career to the next level and make a significant impact, we'd love to hear from you.LSEG is a leading global financial markets infrastructure and data provider. Our purpose is driving financial stability, empowering economies and enabling customers to create sustainable growth.Our purpose is the foundation on which our culture is built. Our values of Integrity, Partnership , Excellence and Change underpin our purpose and set the standard for everything we do, every day. They go to the heart of who we are and guide our decision making and everyday actions.Working with us means that you will be part of a dynamic organisation of 25,000 people across 65 countries. However, we will value your individuality and enable you to bring your true self to work so you can help enrich our diverse workforce.We are proud to be an equal opportunities employer. This means that we do not discriminate on the basis of anyone's race, religion, colour, national origin, gender, sexual orientation, gender identity, gender expression, age, marital status, veteran status, pregnancy or disability, or any other basis protected under applicable law. Conforming with applicable law, we can reasonably accommodate applicants' and employees' religious practices and beliefs, as well as mental health or physical disability needs.You will be part of a collaborative and creative culture where we encourage new ideas. We are committed to sustainability across our global business and we are proud to partner with our customers to help them meet their sustainability objectives. Our charity, the LSEG Foundation provides charitable grants to community groups that help people access economic opportunities and build a secure future with financial independence. Colleagues can get involved through fundraising and volunteering.LSEG offers a range of tailored benefits and support, including healthcare, retirement planning, paid volunteering days and wellbeing initiatives.Please take a moment to read this privacy notice carefully, as it describes what personal information London Stock Exchange Group (LSEG) (we) may hold about you, what it's used for, and how it's obtained, your rights and how to contact us as a data subject.If you are submitting as a Recruitment Agency Partner, it is essential and your responsibility to ensure that candidates applying to LSEG are aware of this privacy notice.
12/07/2026
Full time
Software & Data EngineersApplylocations: London, United Kingdomtime type: Full timeposted on: Posted Todayjob requisition id: R ROLE SUMMARY: FTSE Russell, part of LSEG, powers some of the world's most recognised benchmarks and index solutions. Our technology, data and products help investors measure, manage and capture market opportunities across global financial markets. As we continue to modernise our platforms and scale our capabilities, our engineers, product specialists and data experts play a critical role in delivering innovative solutions that support investment decisions for clients worldwide. WHAT YOU'LL DO: We're looking for a hands-on engineers to help modernise and scale the FTSE platform - a cloud-native, low-latency system that operates at global scale and underpins critical investment products used worldwide.We are recruiting for individual contributors across a varied technical tech. You'll work on challenging problems across software engineering where performance, data quality, and reliability are non-negotiable, while leveraging modern cloud (AWS) and AI-assisted development tooling to accelerate delivery without compromising control.You will be need to have practical use of AI-assisted development tooling to improve engineering productivity. WHAT YOU'LL BRING: Experience developing software using Java, Python, C++, C#, or similar technologies Experience with cloud platforms such as AWS or Azure Knowledge of CI/CD, automation, testing, and modern engineering practices Strong problem-solving and analytical skills Interest in AI and emerging technologies Ability to collaborate effectively across global teams Bachelor's degree in computer science, Engineering, or equivalent practical experience Experience We Value We value curiosity, innovation, and continuous learning, and we're looking for engineers who are excited by emerging technologies and motivated to make a meaningful impact . Join us and help build the technology and data platforms that sit at the heart of global finance. We welcome professionals with experience in one or more of the following areas: Software Engineering Data Engineering Cloud & Platform Engineering Site Reliability Engineering (SRE) AI & Machine Learning DevOps & Automation Architecture & Distributed Systems Analytics & Data PlatformsBring your curiosity, expertise, and ambition-and help build what's next for global investment and market intelligence. ABOUT US: LSEG (London Stock Exchange Group) is more than a diversified global financial markets infrastructure and data business. We are dedicated, open-access partners with a dedication to excellence in delivering the services our customers expect from us. With extensive experience, deep knowledge and worldwide presence across financial markets, we enable businesses and economies around the world to fund innovation, manage risk and create jobs. It's how we've contributed to supporting the financial stability and growth of communities and economies globally for more than 300 years. Through a comprehensive suite of trusted financial market infrastructure services - and our open-access model - we provide the flexibility, stability and trust that enable our customers to pursue their ambitions with confidence and clarity.LSEG is headquartered in the United Kingdom, with significant operations in 70 countries across EMEA, North America, Latin America and Asia Pacific. We employ 25,000 people globally, more than half located in Asia Pacific. LSEG's ticker symbol is LSEG. OUR PEOPLE: People are at the heart of what we do and drive the success of our business. Our culture of connecting, creating opportunity and delivering excellence shape how we think, how we do things and how we help our people fulfil their potential. We embrace diversity and actively seek to attract individuals with unique backgrounds and perspectives. We break down barriers and encourage teamwork, enabling innovation and rapid development of solutions that make a difference. Our workplace generates an enriching and rewarding experience for our people and customers alike. Our vision is to build an inclusive culture in which everyone feels encouraged to fulfil their potential.We know that real personal growth cannot be achieved by simply climbing a career ladder - which is why we encourage and enable a wealth of avenues and interesting opportunities for everyone to broaden and deepen their skills and expertise. As a global organisation spanning 70 countries and one rooted in a culture of growth, opportunity, diversity and innovation, LSEG is a place where everyone can grow, develop and fulfil your potential with meaningful careers Career Stage: Senior Associate London Stock Exchange Group (LSEG) Information: Join us and be part of a team that values innovation, quality, and continuous improvement. If you're ready to take your career to the next level and make a significant impact, we'd love to hear from you.LSEG is a leading global financial markets infrastructure and data provider. Our purpose is driving financial stability, empowering economies and enabling customers to create sustainable growth.Our purpose is the foundation on which our culture is built. Our values of Integrity, Partnership , Excellence and Change underpin our purpose and set the standard for everything we do, every day. They go to the heart of who we are and guide our decision making and everyday actions.Working with us means that you will be part of a dynamic organisation of 25,000 people across 65 countries. However, we will value your individuality and enable you to bring your true self to work so you can help enrich our diverse workforce.We are proud to be an equal opportunities employer. This means that we do not discriminate on the basis of anyone's race, religion, colour, national origin, gender, sexual orientation, gender identity, gender expression, age, marital status, veteran status, pregnancy or disability, or any other basis protected under applicable law. Conforming with applicable law, we can reasonably accommodate applicants' and employees' religious practices and beliefs, as well as mental health or physical disability needs.You will be part of a collaborative and creative culture where we encourage new ideas. We are committed to sustainability across our global business and we are proud to partner with our customers to help them meet their sustainability objectives. Our charity, the LSEG Foundation provides charitable grants to community groups that help people access economic opportunities and build a secure future with financial independence. Colleagues can get involved through fundraising and volunteering.LSEG offers a range of tailored benefits and support, including healthcare, retirement planning, paid volunteering days and wellbeing initiatives.Please take a moment to read this privacy notice carefully, as it describes what personal information London Stock Exchange Group (LSEG) (we) may hold about you, what it's used for, and how it's obtained, your rights and how to contact us as a data subject.If you are submitting as a Recruitment Agency Partner, it is essential and your responsibility to ensure that candidates applying to LSEG are aware of this privacy notice.

Modal Window

  • Home
  • Contact
  • About Us
  • FAQs
  • Terms & Conditions
  • Privacy
  • Employer
  • Post a Job
  • Search Resumes
  • Sign in
  • Job Seeker
  • Find Jobs
  • Create Resume
  • Sign in
  • IT blog
  • Facebook
  • Twitter
  • LinkedIn
  • Youtube
© 2008-2026 IT Job Board