it job board logo
  • Home
  • Find IT Jobs
  • Register CV
  • Career Advice
  • Contact us
  • Employers
    • Register as Employer
    • Pricing Plans
  • Recruiting? Post a job
  • Sign in
  • Sign up
  • Home
  • Find IT Jobs
  • Register CV
  • Career Advice
  • Contact us
  • Employers
    • Register as Employer
    • Pricing Plans
Sorry, that job is no longer available. Here are some results that may be similar to the job you were looking for.

51 jobs found

Email me jobs like this
Refine Search
Current Search
technology architect infrastructure architect capacity performance multi platform lond
Senior Software Engineer (Back End)
TP ICAP Group
The TP ICAP Group is a world leading provider of market infrastructure.Our purpose is to provide clients with access to global financial and commodities markets, improving price discovery, liquidity, and distribution of data, through responsible and innovative solutions.Through our people and technology, we connect clients to superior liquidity and data solutions.The Group is home to a stable of premium brands. Collectively, TP ICAP is the largest interdealer broker in the world by revenue, the number one Energy & Commodities broker in the world, the world's leading provider of OTC data, and an award winning all-to-all trading platform.Founded in London in 1866, the Group operates from more than 60 offices in 27 countries. We are 5,200 people strong. We work as one to achieve our vision of being the world's most trusted, innovative, liquidity and data solutions specialist. Role Overview As a Senior Back-End Engineer, you will design, develop, and maintain robust server-side applications and APIs that support high-volume, data-intensive environments. You will play a key role in shaping technical architecture, ensuring system reliability, and driving best practices in software engineering. This position requires strong problem-solving skills, deep technical expertise, and the ability to collaborate effectively with cross-functional teams. Role Responsibilities Lead by example as a hands on engineer , working closely with Architects, Principal Engineers, and Trading SMEs to design, build, review, test, and deliver mission critical, high performance systems end to end. Own engineering deliverables throughout the full lifecycle, ensuring performance, scalability, resilience, and alignment with engineering best practices. Mentor engineers, elevate code quality standards, and foster a culture of technical excellence. Drive innovation through POCs, technical evaluations, and continuous improvement initiatives. Communicate progress proactively, highlighting risks and removing delivery impediments. Experience / Competences Essential Significant years of hands-on experience building and supporting latency sensitive front office trading systems (OMS, Matching, Execution). Strong background in designing and maintaining distributed, event driven, cloud native applications. Deep understanding of low latency engineering, concurrency, multithreading, and performance optimisation. Comprehensive SDLC experience across design, development, QA, deployment, and production support. Ability to balance rapid delivery with architectural rigour and long term maintainability. Strong relational database design and optimisation skills (MSSQL, MySQL). Proven problem solving capabilities and ability to validate ideas through POCs. Experience building automated testing frameworks for complex distributed systems. Ability to engage effectively with traders, quants, and other business stakeholders.Desired Domain experience in Credit or Fixed Income. Expertise in modern .NET technologies and C# (Java or similar OO languages considered). Experience implementing observability (metrics, tracing, logging) for distributed systems. Experience with CI/CD pipelines, containerisation (Docker), and orchestration (Kubernetes/EKS). Experience designing APIs (REST, GraphQL). Knowledge of distributed messaging and caching technologies (e.g., Solace, Redis, or similar). Understanding of FIX protocol and FIX message handling. Hands-on experience with AWS, microservices, and serverless patterns. Familiarity with React and DAPR (nice to have). Understanding of TDD, BDD, or similar testing methodologies. Role Band & Level: Manager, 6 Company Statement We know that the best innovation happens when diverse people with different perspectives and skills work together in an inclusive atmosphere. That's why we're building a culture where everyone plays a part in making people feel welcome, ready and willing to contribute. TP ICAP Accord - our Employee Network - is a central to this. As well as representing specific groups, TP ICAP Accord helps increase awareness, collaboration, shares best practice, and holds our firm to account for driving continuous cultural improvement. Location UK - 135 Bishopsgate - London Connecting clients, communities and colleagues for sustainable growth TP ICAP connects people, platforms, ideas, and insight across the world's financial, energy and commodities markets. As a global leader in market infrastructure and data-led solutions, we enhance market access, increase efficiencies, and unlock possibilities. Work with us Joining TP ICAP puts you at the heart of markets that matter.You'll have the freedom to innovate and act on your initiative. We'll train you and build your abilities in your specialist area, so that you can become an expert in your field. And all within a connected network that's there to set you up for success.TP ICAP Group is a collection of premium brands each with a distinct, client-focused offering. Underpinning and connecting these client-facing brands is the financial security, operational strength and know-how we have as a Group.Connections are at the heart of what we do. We combine our people's know-how with the latest technology to improve price discovery, trade execution and liquidity flow.Connections create strength. Through them, we help our clients to manage risk, realise investment strategies and expand the scope for growth.And connections act as a catalyst. Sparking richer solutions for our clients to break new ground, modernising markets for future performance, and creating dynamic careers for our people. Our capacity to connect builds trust, supports communities and gives us the power to anticipate and respond to change, whatever direction the world takes. It's what makes TP ICAP a mainstay in the global markets, now and in the future.TP ICAP. We connect.
27/07/2026
Full time
The TP ICAP Group is a world leading provider of market infrastructure.Our purpose is to provide clients with access to global financial and commodities markets, improving price discovery, liquidity, and distribution of data, through responsible and innovative solutions.Through our people and technology, we connect clients to superior liquidity and data solutions.The Group is home to a stable of premium brands. Collectively, TP ICAP is the largest interdealer broker in the world by revenue, the number one Energy & Commodities broker in the world, the world's leading provider of OTC data, and an award winning all-to-all trading platform.Founded in London in 1866, the Group operates from more than 60 offices in 27 countries. We are 5,200 people strong. We work as one to achieve our vision of being the world's most trusted, innovative, liquidity and data solutions specialist. Role Overview As a Senior Back-End Engineer, you will design, develop, and maintain robust server-side applications and APIs that support high-volume, data-intensive environments. You will play a key role in shaping technical architecture, ensuring system reliability, and driving best practices in software engineering. This position requires strong problem-solving skills, deep technical expertise, and the ability to collaborate effectively with cross-functional teams. Role Responsibilities Lead by example as a hands on engineer , working closely with Architects, Principal Engineers, and Trading SMEs to design, build, review, test, and deliver mission critical, high performance systems end to end. Own engineering deliverables throughout the full lifecycle, ensuring performance, scalability, resilience, and alignment with engineering best practices. Mentor engineers, elevate code quality standards, and foster a culture of technical excellence. Drive innovation through POCs, technical evaluations, and continuous improvement initiatives. Communicate progress proactively, highlighting risks and removing delivery impediments. Experience / Competences Essential Significant years of hands-on experience building and supporting latency sensitive front office trading systems (OMS, Matching, Execution). Strong background in designing and maintaining distributed, event driven, cloud native applications. Deep understanding of low latency engineering, concurrency, multithreading, and performance optimisation. Comprehensive SDLC experience across design, development, QA, deployment, and production support. Ability to balance rapid delivery with architectural rigour and long term maintainability. Strong relational database design and optimisation skills (MSSQL, MySQL). Proven problem solving capabilities and ability to validate ideas through POCs. Experience building automated testing frameworks for complex distributed systems. Ability to engage effectively with traders, quants, and other business stakeholders.Desired Domain experience in Credit or Fixed Income. Expertise in modern .NET technologies and C# (Java or similar OO languages considered). Experience implementing observability (metrics, tracing, logging) for distributed systems. Experience with CI/CD pipelines, containerisation (Docker), and orchestration (Kubernetes/EKS). Experience designing APIs (REST, GraphQL). Knowledge of distributed messaging and caching technologies (e.g., Solace, Redis, or similar). Understanding of FIX protocol and FIX message handling. Hands-on experience with AWS, microservices, and serverless patterns. Familiarity with React and DAPR (nice to have). Understanding of TDD, BDD, or similar testing methodologies. Role Band & Level: Manager, 6 Company Statement We know that the best innovation happens when diverse people with different perspectives and skills work together in an inclusive atmosphere. That's why we're building a culture where everyone plays a part in making people feel welcome, ready and willing to contribute. TP ICAP Accord - our Employee Network - is a central to this. As well as representing specific groups, TP ICAP Accord helps increase awareness, collaboration, shares best practice, and holds our firm to account for driving continuous cultural improvement. Location UK - 135 Bishopsgate - London Connecting clients, communities and colleagues for sustainable growth TP ICAP connects people, platforms, ideas, and insight across the world's financial, energy and commodities markets. As a global leader in market infrastructure and data-led solutions, we enhance market access, increase efficiencies, and unlock possibilities. Work with us Joining TP ICAP puts you at the heart of markets that matter.You'll have the freedom to innovate and act on your initiative. We'll train you and build your abilities in your specialist area, so that you can become an expert in your field. And all within a connected network that's there to set you up for success.TP ICAP Group is a collection of premium brands each with a distinct, client-focused offering. Underpinning and connecting these client-facing brands is the financial security, operational strength and know-how we have as a Group.Connections are at the heart of what we do. We combine our people's know-how with the latest technology to improve price discovery, trade execution and liquidity flow.Connections create strength. Through them, we help our clients to manage risk, realise investment strategies and expand the scope for growth.And connections act as a catalyst. Sparking richer solutions for our clients to break new ground, modernising markets for future performance, and creating dynamic careers for our people. Our capacity to connect builds trust, supports communities and gives us the power to anticipate and respond to change, whatever direction the world takes. It's what makes TP ICAP a mainstay in the global markets, now and in the future.TP ICAP. We connect.
SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom
Hitachids
We're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people-centric and here to power good.Every day, we future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our companyand customers from what's now to what's next.We make it happen through the power of acceleration. Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us. Job Description Mandatory Skills: Observability, Resiliency, Service Management, Reliability, Performance engineering, Scalability, release management, Cloud cost management. Role Description Skills: ROLE PURPOSE Lead the Site Reliability Engineering practice, driving the transformation from reactive operations to proactive, engineering-led reliability. Own the definition and enforcement of non-functional requirements (NFRs) using FMEA-based resiliency frameworks, and champion observability, self-healing automation, automated incident management, and database operations automation. Ensure systems are resilient, performant, cost-optimised, and continuously improving. KEY RESPONSIBILITIES Define and enforce non-functional requirements (NFRs) for performance, scalability, availability, fault tolerance, and cost efficiency using FMEA-based failure analysis Design and implement self-healing automation for known failure patterns, reducing human intervention and on-call burden by 50%+ Build comprehensive observability stacks (metrics, logs, traces) with ML-driven anomaly detection and AIOps capabilities Lead automated incident management: detection, triage, escalation, remediation, and post-incident review automation Drive DB automation: automated provisioning, release management (UK focus), backup/restore, and operational request workflows for all database operations Define and track SLIs, SLOs, and error budgets across all critical services, using them to balance reliability with feature velocity Conduct chaos engineering exercises and game days to validate resiliency and uncover hidden failure modes Mentor 2 SRE Engineers, establish engineering standards, and build a culture of reliability and continuous improvement Collaborate with Platform Engineering and Cloud teams to embed reliability into infrastructure and deployment pipelines TECHNICAL SKILLS & EXPERTISE Expert-level observability: Prometheus, Grafana, ELK/OpenSearch, Jaeger/Zipkin, Datadog, or Dynatrace Strong experience with AIOps and ML-driven monitoring: PagerDuty, Moogsoft, BigPanda, or custom ML pipelines Deep knowledge of FMEA, fault tree analysis, and chaos engineering tools (Gremlin, LitmusChaos, Chaos Monkey) Database automation: strong SQL skills plus experience with automated DB provisioning, migration tools (Liquibase, Flyway), and DB release pipelines Proficiency in automation and scripting: Python, Go, Bash, with experience building self-healing runbooks Infrastructure knowledge: Kubernetes, cloud platforms (AWS/Azure/GCP), networking, and storage systems CI/CD and release engineering: Jenkins, GitLab CI, Spinnaker, ArgoCD for integrated DB and application releases Cost management: experience with FinOps principles, resource optimisation, and cloud spend analysis SOFT SKILLS & COMPETENCIES Strong leadership and mentoring ability - coaches and develops junior engineers Excellent stakeholder management and communication skills across all levels Strategic thinker who balances technical depth with business outcomes Proven ability to drive change, influence without authority, and build consensus Strong analytical and problem-solving mindset with attention to detail Ability to manage competing priorities across multiple workstreams simultaneously QUALIFICATIONS & EXPERIENCE 7+ years in SRE, DevOps, or production engineering with 3+ years in a senior or lead capacity Proven track record of improving availability, reducing MTTR, and implementing self-healing at scale Experience managing or automating database operations in enterprise environments Relevant certifications preferred: CKA, AWS DevOps Professional, Azure DevOps Expert, SRE Foundation Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience) DESIRABLE / NICE TO HAVE Published work or conference talks on SRE, observability, or chaos engineering Experience with service mesh (Istio) and distributed tracing at scale Background in financial services or regulated industry SRE practices About Us We're a global, team of innovators. Together, we harness engineering excellence and passion to co-create meaningful solutions to complex challenges. We turn organizations into data-driven leaders that can make a positive impact on their industries and society. If you believe that innovation can bring a better tomorrow closer to today, this is the place for you. Fostering innovation through diverse perspectives Hitachi is a global company operating across a wide range of industries and regions. One of the things that sets Hitachi apart is the diversity of our business and people, which drives our innovation and growth. We are committed to building an inclusive culture based on mutual respect and merit-based systems. We believe that when people feel valued, heard, and safe to express themselves, they do their best work. How we look after you We help take care of your today and tomorrow with industry-leading benefits, support, and services that look after your holistic health and wellbeing. We're also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We're always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you'll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with. We're proud to say we're an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
27/07/2026
Full time
We're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people-centric and here to power good.Every day, we future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our companyand customers from what's now to what's next.We make it happen through the power of acceleration. Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us. Job Description Mandatory Skills: Observability, Resiliency, Service Management, Reliability, Performance engineering, Scalability, release management, Cloud cost management. Role Description Skills: ROLE PURPOSE Lead the Site Reliability Engineering practice, driving the transformation from reactive operations to proactive, engineering-led reliability. Own the definition and enforcement of non-functional requirements (NFRs) using FMEA-based resiliency frameworks, and champion observability, self-healing automation, automated incident management, and database operations automation. Ensure systems are resilient, performant, cost-optimised, and continuously improving. KEY RESPONSIBILITIES Define and enforce non-functional requirements (NFRs) for performance, scalability, availability, fault tolerance, and cost efficiency using FMEA-based failure analysis Design and implement self-healing automation for known failure patterns, reducing human intervention and on-call burden by 50%+ Build comprehensive observability stacks (metrics, logs, traces) with ML-driven anomaly detection and AIOps capabilities Lead automated incident management: detection, triage, escalation, remediation, and post-incident review automation Drive DB automation: automated provisioning, release management (UK focus), backup/restore, and operational request workflows for all database operations Define and track SLIs, SLOs, and error budgets across all critical services, using them to balance reliability with feature velocity Conduct chaos engineering exercises and game days to validate resiliency and uncover hidden failure modes Mentor 2 SRE Engineers, establish engineering standards, and build a culture of reliability and continuous improvement Collaborate with Platform Engineering and Cloud teams to embed reliability into infrastructure and deployment pipelines TECHNICAL SKILLS & EXPERTISE Expert-level observability: Prometheus, Grafana, ELK/OpenSearch, Jaeger/Zipkin, Datadog, or Dynatrace Strong experience with AIOps and ML-driven monitoring: PagerDuty, Moogsoft, BigPanda, or custom ML pipelines Deep knowledge of FMEA, fault tree analysis, and chaos engineering tools (Gremlin, LitmusChaos, Chaos Monkey) Database automation: strong SQL skills plus experience with automated DB provisioning, migration tools (Liquibase, Flyway), and DB release pipelines Proficiency in automation and scripting: Python, Go, Bash, with experience building self-healing runbooks Infrastructure knowledge: Kubernetes, cloud platforms (AWS/Azure/GCP), networking, and storage systems CI/CD and release engineering: Jenkins, GitLab CI, Spinnaker, ArgoCD for integrated DB and application releases Cost management: experience with FinOps principles, resource optimisation, and cloud spend analysis SOFT SKILLS & COMPETENCIES Strong leadership and mentoring ability - coaches and develops junior engineers Excellent stakeholder management and communication skills across all levels Strategic thinker who balances technical depth with business outcomes Proven ability to drive change, influence without authority, and build consensus Strong analytical and problem-solving mindset with attention to detail Ability to manage competing priorities across multiple workstreams simultaneously QUALIFICATIONS & EXPERIENCE 7+ years in SRE, DevOps, or production engineering with 3+ years in a senior or lead capacity Proven track record of improving availability, reducing MTTR, and implementing self-healing at scale Experience managing or automating database operations in enterprise environments Relevant certifications preferred: CKA, AWS DevOps Professional, Azure DevOps Expert, SRE Foundation Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience) DESIRABLE / NICE TO HAVE Published work or conference talks on SRE, observability, or chaos engineering Experience with service mesh (Istio) and distributed tracing at scale Background in financial services or regulated industry SRE practices About Us We're a global, team of innovators. Together, we harness engineering excellence and passion to co-create meaningful solutions to complex challenges. We turn organizations into data-driven leaders that can make a positive impact on their industries and society. If you believe that innovation can bring a better tomorrow closer to today, this is the place for you. Fostering innovation through diverse perspectives Hitachi is a global company operating across a wide range of industries and regions. One of the things that sets Hitachi apart is the diversity of our business and people, which drives our innovation and growth. We are committed to building an inclusive culture based on mutual respect and merit-based systems. We believe that when people feel valued, heard, and safe to express themselves, they do their best work. How we look after you We help take care of your today and tomorrow with industry-leading benefits, support, and services that look after your holistic health and wellbeing. We're also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We're always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you'll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with. We're proud to say we're an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
Programme Manager - Energy & Commodities
TP ICAP Group Services Ltd City Of Westminster, London
Company Overview The TP ICAP Group is a world leading provider of market infrastructure. Our purpose is to provide clients with access to global financial and commodities markets, improving price discovery, liquidity, and distribution of data, through responsible and innovative solutions. Through our people and technology, we connect clients to superior liquidity and data solutions. The Group is home to a stable of premium brands. Collectively, TP ICAP is the largest interdealer broker in the world by revenue, the number one Energy & Commodities broker in the world, the world's leading provider of OTC data, and an award winning all-to-all trading platform. Founded in London in 1866, the Group operates from more than 60 offices in 27 countries. We are 5,200 people strong. We work as one to achieve our vision of being the world's most trusted, innovative, liquidity and data solutions specialist. Role Overview The Programme Manager is responsible for leading the technology portfolio and coordinating the end to end delivery of complex, multi workstream programmes that support organisational transformation across technology and the Energy & Commodities business. The role includes accountability for programme governance, cross functional delivery, stakeholder management, risk and issue management, and alignment of programme outcomes with strategic objectives. This position requires strong leadership, excellent communication, and the ability to manage diverse teams and dependencies across global locations. Role Responsibilities Partner with the Business Sponsor to ensure programme initiatives align with strategic portfolio objectives and business outcomes Lead end-to-end delivery of portfolio and programme objectives across multiple projects and workstreams Define, maintain and communicate the programme roadmap, ensuring alignment with organisational priorities Deliver change through the TP ICAP Change Management framework, maintaining appropriate controls and delivery standards Manage portfolio budgets and financial performance, ensuring effective tracking and delivery within agreed constraints Establish and maintain governance structures in line with TP ICAP corporate standards, providing appropriate oversight, reporting and decision-making Identify, manage and mitigate portfolio-level risks, issues and dependencies, ensuring impacts are understood and communicated Build and maintain strong stakeholder relationships across business and technology teams, providing clear communication on progress, decisions and outcomes Oversee resource planning and allocation, including Project Managers and Business Analysts, to maximise delivery effectiveness Agree Business Analyst deliverables with Business Analysis Management and ensure quality outcomes Support change demand, business case development and programme/project initiation activities Facilitate workshops, working groups and cross-functional forums to drive alignment, decision-making and delivery Coach, mentor and motivate team members ensuring accountability, collaboration and high performance Provide concise and timely reporting to senior stakeholders and governance forums Undertake additional responsibilities as required to support evolving business needs Manage Energy and Commodities vendor relationships from a technology perspective Fulfil any additional / ad hoc duties as required to meet the needs of the business Experience / Competences Essential Proven track record of delivering complex programmes and strategic change initiatives within organisations of comparable scale, complexity and regulatory environment to TP ICAP, ideally within Commodities, Energy or Financial Services. Demonstrated experience managing geographically dispersed delivery teams and coordinating outcomes across multiple locations and functions. Strong technical acumen with the ability to understand complex architectures, navigate technical challenges and support effective solution delivery. Extensive knowledge of programme delivery methodologies, change management disciplines and governance frameworks, ideally SAFe Excellent communication and stakeholder management skills, with the ability to influence and engage senior leaders and executive stakeholders. Highly organised, with a proven ability to manage competing priorities, complex dependencies and challenging delivery timelines. Strong analytical and problem-solving capabilities, with the ability to simplify complex issues, drive clarity and support informed decision-making. Possesses professional experience in programme, project or change management. Experience managing and controlling programme budgets, forecasts and financial performance. Resilient, adaptable and delivery-focused, with a collaborative and service-oriented approach in fast-paced environments. Proven ability to lead, develop and motivate high-performing teams, showing accountability, continuous improvement and successful delivery outcomes. Results-driven with a strong sense of ownership, accountability and a focus on delivering measurable business value. Adaptable, resilient and calm under pressure, maintaining focus and delivery momentum through periods of change and uncertainty. Self motivated, conscientious and goal-oriented, with a strong commitment to high-quality delivery and continuous improvement. Desired Desired Degree qualified or possessing an equivalent combination of professional experience and relevant qualifications, e.g., PMI, SAFe Applied engineering capability to deliver business solutions throughout the development lifecycle Demonstrated ability to lead large scale transformation, regulatory and business change programmes from inception through to successful delivery Strong leadership and stakeholder management capabilities, with the ability to influence, challenge and gain commitment from senior executives and cross functional teams Experienced in working across both business and technology teams, translating requirements into actionable outcomes and driving alignment Role Band & Level: Manager 7 Location UK - 3rd Floor Verde Building - London Company Statement We know that the best innovation happens when diverse people with different perspectives and skills work together in an inclusive atmosphere. That's why we're building a culture where everyone plays a part in making people feel welcome, ready and willing to contribute. TP ICAP Accord - our Employee Network - is a central to this. As well as representing specific groups, TP ICAP Accord helps increase awareness, collaboration, shares best practice, and holds our firm to account for driving continuous cultural improvement. Work with us Joining TP ICAP puts you at the heart of markets that matter. You'll have the freedom to innovate and act on your initiative. We'll train you and build your abilities in your specialist area, so that you can become an expert in your field. And all within a connected network that's there to set you up for success. More about us TP ICAP Group is a collection of premium brands each with a distinct, client focused offering. Underpinning and connecting these client facing brands is the financial security, operational strength and know how we have as a Group. Connections are at the heart of what we do. We combine our people's know how with the latest technology to improve price discovery, trade execution and liquidity flow. Connections create strength. Through them, we help our clients to manage risk, realise investment strategies and expand the scope for growth. And connections act as a catalyst. Sparking richer solutions for our clients to break new ground, modernising markets for future performance, and creating dynamic careers for our people. Our capacity to connect builds trust, supports communities and gives us the power to anticipate and respond to change, whatever direction the world takes. It's what makes TP ICAP a mainstay in the global markets, now and in the future. TP ICAP. We connect.
26/07/2026
Full time
Company Overview The TP ICAP Group is a world leading provider of market infrastructure. Our purpose is to provide clients with access to global financial and commodities markets, improving price discovery, liquidity, and distribution of data, through responsible and innovative solutions. Through our people and technology, we connect clients to superior liquidity and data solutions. The Group is home to a stable of premium brands. Collectively, TP ICAP is the largest interdealer broker in the world by revenue, the number one Energy & Commodities broker in the world, the world's leading provider of OTC data, and an award winning all-to-all trading platform. Founded in London in 1866, the Group operates from more than 60 offices in 27 countries. We are 5,200 people strong. We work as one to achieve our vision of being the world's most trusted, innovative, liquidity and data solutions specialist. Role Overview The Programme Manager is responsible for leading the technology portfolio and coordinating the end to end delivery of complex, multi workstream programmes that support organisational transformation across technology and the Energy & Commodities business. The role includes accountability for programme governance, cross functional delivery, stakeholder management, risk and issue management, and alignment of programme outcomes with strategic objectives. This position requires strong leadership, excellent communication, and the ability to manage diverse teams and dependencies across global locations. Role Responsibilities Partner with the Business Sponsor to ensure programme initiatives align with strategic portfolio objectives and business outcomes Lead end-to-end delivery of portfolio and programme objectives across multiple projects and workstreams Define, maintain and communicate the programme roadmap, ensuring alignment with organisational priorities Deliver change through the TP ICAP Change Management framework, maintaining appropriate controls and delivery standards Manage portfolio budgets and financial performance, ensuring effective tracking and delivery within agreed constraints Establish and maintain governance structures in line with TP ICAP corporate standards, providing appropriate oversight, reporting and decision-making Identify, manage and mitigate portfolio-level risks, issues and dependencies, ensuring impacts are understood and communicated Build and maintain strong stakeholder relationships across business and technology teams, providing clear communication on progress, decisions and outcomes Oversee resource planning and allocation, including Project Managers and Business Analysts, to maximise delivery effectiveness Agree Business Analyst deliverables with Business Analysis Management and ensure quality outcomes Support change demand, business case development and programme/project initiation activities Facilitate workshops, working groups and cross-functional forums to drive alignment, decision-making and delivery Coach, mentor and motivate team members ensuring accountability, collaboration and high performance Provide concise and timely reporting to senior stakeholders and governance forums Undertake additional responsibilities as required to support evolving business needs Manage Energy and Commodities vendor relationships from a technology perspective Fulfil any additional / ad hoc duties as required to meet the needs of the business Experience / Competences Essential Proven track record of delivering complex programmes and strategic change initiatives within organisations of comparable scale, complexity and regulatory environment to TP ICAP, ideally within Commodities, Energy or Financial Services. Demonstrated experience managing geographically dispersed delivery teams and coordinating outcomes across multiple locations and functions. Strong technical acumen with the ability to understand complex architectures, navigate technical challenges and support effective solution delivery. Extensive knowledge of programme delivery methodologies, change management disciplines and governance frameworks, ideally SAFe Excellent communication and stakeholder management skills, with the ability to influence and engage senior leaders and executive stakeholders. Highly organised, with a proven ability to manage competing priorities, complex dependencies and challenging delivery timelines. Strong analytical and problem-solving capabilities, with the ability to simplify complex issues, drive clarity and support informed decision-making. Possesses professional experience in programme, project or change management. Experience managing and controlling programme budgets, forecasts and financial performance. Resilient, adaptable and delivery-focused, with a collaborative and service-oriented approach in fast-paced environments. Proven ability to lead, develop and motivate high-performing teams, showing accountability, continuous improvement and successful delivery outcomes. Results-driven with a strong sense of ownership, accountability and a focus on delivering measurable business value. Adaptable, resilient and calm under pressure, maintaining focus and delivery momentum through periods of change and uncertainty. Self motivated, conscientious and goal-oriented, with a strong commitment to high-quality delivery and continuous improvement. Desired Desired Degree qualified or possessing an equivalent combination of professional experience and relevant qualifications, e.g., PMI, SAFe Applied engineering capability to deliver business solutions throughout the development lifecycle Demonstrated ability to lead large scale transformation, regulatory and business change programmes from inception through to successful delivery Strong leadership and stakeholder management capabilities, with the ability to influence, challenge and gain commitment from senior executives and cross functional teams Experienced in working across both business and technology teams, translating requirements into actionable outcomes and driving alignment Role Band & Level: Manager 7 Location UK - 3rd Floor Verde Building - London Company Statement We know that the best innovation happens when diverse people with different perspectives and skills work together in an inclusive atmosphere. That's why we're building a culture where everyone plays a part in making people feel welcome, ready and willing to contribute. TP ICAP Accord - our Employee Network - is a central to this. As well as representing specific groups, TP ICAP Accord helps increase awareness, collaboration, shares best practice, and holds our firm to account for driving continuous cultural improvement. Work with us Joining TP ICAP puts you at the heart of markets that matter. You'll have the freedom to innovate and act on your initiative. We'll train you and build your abilities in your specialist area, so that you can become an expert in your field. And all within a connected network that's there to set you up for success. More about us TP ICAP Group is a collection of premium brands each with a distinct, client focused offering. Underpinning and connecting these client facing brands is the financial security, operational strength and know how we have as a Group. Connections are at the heart of what we do. We combine our people's know how with the latest technology to improve price discovery, trade execution and liquidity flow. Connections create strength. Through them, we help our clients to manage risk, realise investment strategies and expand the scope for growth. And connections act as a catalyst. Sparking richer solutions for our clients to break new ground, modernising markets for future performance, and creating dynamic careers for our people. Our capacity to connect builds trust, supports communities and gives us the power to anticipate and respond to change, whatever direction the world takes. It's what makes TP ICAP a mainstay in the global markets, now and in the future. TP ICAP. We connect.
SRE Architect (68019)
Hitachi Automotive Systems Americas, Inc.
FunctionCloud & Data EngineeringOur CompanyWe're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people-centric and here to power good. Every day, we future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our company and customers from what's now to what's next. We make it happen through the power of acceleration.Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us.Job descriptionMandatory Skills:Observability, Resiliency, Service Management, Reliability, Performance engineering, Scalability, release management, Cloud cost management.Role Description Skills:ROLE PURPOSELead the Site Reliability Engineering practice, driving the transformation from reactive operations to proactive, engineering-led reliability. Own the definition and enforcement of non-functional requirements (NFRs) using FMEA-based resiliency frameworks, and champion observability, self-healing automation, automated incident management, and database operations automation. Ensure systems are resilient, performant, cost-optimised, and continuously improving.KEY RESPONSIBILITIESDefine and enforce non-functional requirements (NFRs) for performance, scalability, availability, fault tolerance, and cost efficiency using FMEA-based failure analysisDesign and implement self-healing automation for known failure patterns, reducing human intervention and on-call burden by 50%+Build comprehensive observability stacks (metrics, logs, traces) with ML-driven anomaly detection and AIOps capabilitiesLead automated incident management: detection, triage, escalation, remediation, and post-incident review automationDrive DB automation: automated provisioning, release management (UK focus), backup/restore, and operational request workflows for all database operationsDefine and track SLIs, SLOs, and error budgets across all critical services, using them to balance reliability with feature velocityConduct chaos engineering exercises and game days to validate resiliency and uncover hidden failure modesMentor 2 SRE Engineers, establish engineering standards, and build a culture of reliability and continuous improvementCollaborate with Platform Engineering and Cloud teams to embed reliability into infrastructure and deployment pipelinesTECHNICAL SKILLS & EXPERTISEExpert-level observability: Prometheus, Grafana, ELK/OpenSearch, Jaeger/Zipkin, Datadog, or DynatraceStrong experience with AIOps and ML-driven monitoring: PagerDuty, Moogsoft, BigPanda, or custom ML pipelinesDeep knowledge of FMEA, fault tree analysis, and chaos engineering tools (Gremlin, LitmusChaos, Chaos Monkey)Database automation: strong SQL skills plus experience with automated DB provisioning, migration tools (Liquibase, Flyway), and DB release pipelinesProficiency in automation and scripting: Python, Go, Bash, with experience building self-healing runbooksInfrastructure knowledge: Kubernetes, cloud platforms (AWS/Azure/GCP), networking, and storage systemsCI/CD and release engineering: Jenkins, GitLab CI, Spinnaker, ArgoCD for integrated DB and application releasesCost management: experience with FinOps principles, resource optimisation, and cloud spend analysisSOFT SKILLS & COMPETENCIESStrong leadership and mentoring ability - coaches and develops junior engineersExcellent stakeholder management and communication skills across all levelsStrategic thinker who balances technical depth with business outcomesProven ability to drive change, influence without authority, and build consensusStrong analytical and problem-solving mindset with attention to detailAbility to manage competing priorities across multiple workstreams simultaneouslyQUALIFICATIONS & EXPERIENCE7+ years in SRE, DevOps, or production engineering with 3+ years in a senior or lead capacityProven track record of improving availability, reducing MTTR, and implementing self-healing at scaleExperience managing or automating database operations in enterprise environmentsRelevant certifications preferred: CKA, AWS DevOps Professional, Azure DevOps Expert, SRE FoundationBachelor's degree in Computer Science, Engineering, or related field (or equivalent experience)DESIRABLE / NICE TO HAVEPublished work or conference talks on SRE, observability, or chaos engineeringExperience with service mesh (Istio) and distributed tracing at scaleBackground in financial services or regulated industry SRE practicesAbout usWe're a global, team of innovators. Together, we harness engineering excellence and passion to co-create meaningful solutions to complex challenges. We turn organizations into data-driven leaders that can make a positive impact on their industries and society. If you believe that innovation can bring a better tomorrow closer to today, this is the place for you.Fostering innovation through diverse perspectivesHitachi is a global company operating across a wide range of industries and regions. One of the things that sets Hitachi apart is the diversity of our business and people, which drives our innovation and growth.We are committed to building an inclusive culture based on mutual respect and merit-based systems. We believe that when people feel valued, heard, and safe to express themselves, they do their best work.How we look after youWe help take care of your today and tomorrow with industry-leading benefits, support, and services that look after your holistic health and wellbeing. We're also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We're always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you'll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with.We're proud to say we're an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
26/07/2026
Full time
FunctionCloud & Data EngineeringOur CompanyWe're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people-centric and here to power good. Every day, we future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our company and customers from what's now to what's next. We make it happen through the power of acceleration.Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us.Job descriptionMandatory Skills:Observability, Resiliency, Service Management, Reliability, Performance engineering, Scalability, release management, Cloud cost management.Role Description Skills:ROLE PURPOSELead the Site Reliability Engineering practice, driving the transformation from reactive operations to proactive, engineering-led reliability. Own the definition and enforcement of non-functional requirements (NFRs) using FMEA-based resiliency frameworks, and champion observability, self-healing automation, automated incident management, and database operations automation. Ensure systems are resilient, performant, cost-optimised, and continuously improving.KEY RESPONSIBILITIESDefine and enforce non-functional requirements (NFRs) for performance, scalability, availability, fault tolerance, and cost efficiency using FMEA-based failure analysisDesign and implement self-healing automation for known failure patterns, reducing human intervention and on-call burden by 50%+Build comprehensive observability stacks (metrics, logs, traces) with ML-driven anomaly detection and AIOps capabilitiesLead automated incident management: detection, triage, escalation, remediation, and post-incident review automationDrive DB automation: automated provisioning, release management (UK focus), backup/restore, and operational request workflows for all database operationsDefine and track SLIs, SLOs, and error budgets across all critical services, using them to balance reliability with feature velocityConduct chaos engineering exercises and game days to validate resiliency and uncover hidden failure modesMentor 2 SRE Engineers, establish engineering standards, and build a culture of reliability and continuous improvementCollaborate with Platform Engineering and Cloud teams to embed reliability into infrastructure and deployment pipelinesTECHNICAL SKILLS & EXPERTISEExpert-level observability: Prometheus, Grafana, ELK/OpenSearch, Jaeger/Zipkin, Datadog, or DynatraceStrong experience with AIOps and ML-driven monitoring: PagerDuty, Moogsoft, BigPanda, or custom ML pipelinesDeep knowledge of FMEA, fault tree analysis, and chaos engineering tools (Gremlin, LitmusChaos, Chaos Monkey)Database automation: strong SQL skills plus experience with automated DB provisioning, migration tools (Liquibase, Flyway), and DB release pipelinesProficiency in automation and scripting: Python, Go, Bash, with experience building self-healing runbooksInfrastructure knowledge: Kubernetes, cloud platforms (AWS/Azure/GCP), networking, and storage systemsCI/CD and release engineering: Jenkins, GitLab CI, Spinnaker, ArgoCD for integrated DB and application releasesCost management: experience with FinOps principles, resource optimisation, and cloud spend analysisSOFT SKILLS & COMPETENCIESStrong leadership and mentoring ability - coaches and develops junior engineersExcellent stakeholder management and communication skills across all levelsStrategic thinker who balances technical depth with business outcomesProven ability to drive change, influence without authority, and build consensusStrong analytical and problem-solving mindset with attention to detailAbility to manage competing priorities across multiple workstreams simultaneouslyQUALIFICATIONS & EXPERIENCE7+ years in SRE, DevOps, or production engineering with 3+ years in a senior or lead capacityProven track record of improving availability, reducing MTTR, and implementing self-healing at scaleExperience managing or automating database operations in enterprise environmentsRelevant certifications preferred: CKA, AWS DevOps Professional, Azure DevOps Expert, SRE FoundationBachelor's degree in Computer Science, Engineering, or related field (or equivalent experience)DESIRABLE / NICE TO HAVEPublished work or conference talks on SRE, observability, or chaos engineeringExperience with service mesh (Istio) and distributed tracing at scaleBackground in financial services or regulated industry SRE practicesAbout usWe're a global, team of innovators. Together, we harness engineering excellence and passion to co-create meaningful solutions to complex challenges. We turn organizations into data-driven leaders that can make a positive impact on their industries and society. If you believe that innovation can bring a better tomorrow closer to today, this is the place for you.Fostering innovation through diverse perspectivesHitachi is a global company operating across a wide range of industries and regions. One of the things that sets Hitachi apart is the diversity of our business and people, which drives our innovation and growth.We are committed to building an inclusive culture based on mutual respect and merit-based systems. We believe that when people feel valued, heard, and safe to express themselves, they do their best work.How we look after youWe help take care of your today and tomorrow with industry-leading benefits, support, and services that look after your holistic health and wellbeing. We're also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We're always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you'll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with.We're proud to say we're an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
Software Development Manager
Club L London Stretford, Lancashire
About Us Club L London is the next-generation online fashion retailer for the forward-thinking woman. Conceptualised and crafted in-house and abroad, we specialise in accessible luxury and designs of unrivalled quality that flatter all figures. From prom to occasion, maternity, bridal, and beyond, we deliver an elevated shopping experience that connects our global community of trend-setting consumers, influencers, and content creators with fresh collections dropping weekly. Our engineering team builds and runs the platforms behind Club L London and Lavish Alice across multiple markets - and we're scaling fast. The Role We are looking for a Software Development Manager to lead and grow our engineering team. This is a hands on leadership role that blends people management, delivery ownership, and technical direction. You'll be responsible for the developers who build and maintain our e commerce platforms, integrations, and data systems - setting the standards, shaping the roadmap, and making sure we ship reliable, scalable software at pace. You'll manage a team of full stack and backend engineers, own delivery end to end, and remain close enough to the code to make sound architectural calls and unblock your team. You'll report to the CTO and work across a broad, integration dense stack spanning Shopify Plus, Node.js services, serverless, and a modern GCP/BigQuery data platform. Key Responsibilities Team Leadership & People Management Lead, mentor, and line manage a team of full stack and backend engineers, owning their growth, performance, and development. Run regular 1:1s, set clear objectives, and give consistent, actionable feedback. Drive hiring - define roles, interview, and build a high performing, collaborative engineering culture. Manage team capacity, balance workloads, and protect focus time so the team can deliver sustainably. Delivery & Programme Management Own delivery of the engineering roadmap, translating business priorities into a clear, achievable plan. Run agile ceremonies (planning, stand ups, retrospectives) and keep work moving through the pipeline. Balance BAU, incident response, and project work, managing dependencies, risk, and trade offs. Report progress, blockers, and delivery forecasts clearly to the CTO and wider stakeholders. Technical Leadership & Architecture Guide architecture across e commerce platforms, internal tools, and data systems, ensuring reliability, performance, scalability, and security. Set direction for backend services and APIs (Node.js, REST/GraphQL), serverless functions, and microservices (AWS Lambda or equivalent). Oversee data infrastructure and pipelines across platforms such as BigQuery, MongoDB, or PostgreSQL, ensuring efficient data models and performant queries. Establish and enforce engineering standards - code review culture, testing, CI/CD, and modern development workflows. Platform Integrations & Automation Own the integration landscape connecting e commerce platforms, ERP, marketing, fulfilment, and third party services. Identify opportunities for automation and system improvement across the business, and turn them into delivered outcomes. Ensure integrations are robust, well tested, observable, and resilient as the business scales. Manage technical relationships with key vendors and platform partners. Cross Functional Collaboration & Stakeholder Management Work closely with e commerce, design, content, operations, data, and finance teams to deliver reliable systems and features. Translate business needs into technical delivery, and communicate technical concepts clearly to non technical stakeholders. Champion engineering best practice across the wider tech function. Hands On Contribution Stay technically credible - contribute to code, architecture, code reviews, and complex problem solving where it adds the most value. Lead by example, setting the bar for quality, pragmatism, and delivery. More About You 5+ years in software engineering, with 2+ years leading or line managing engineers, ideally within retail, e commerce, or high traffic digital environments. Proven people leadership - mentoring, performance management, and hiring high performing teams. Strong hands on background with Node.js or similar backend technologies in production. Experience designing and building APIs (REST or GraphQL) and integrating with third party systems. Hands on experience with serverless architectures (AWS Lambda or equivalent). Experience with databases and data platforms such as MongoDB, BigQuery, PostgreSQL, or similar. Solid grounding in CI/CD pipelines, version control (Git), and modern development workflows. A track record of delivering software in an agile environment - planning, prioritising, and shipping reliably. Comfortable balancing hands on technical work with the demands of leading a team. Nice To Have Experience with Shopify (Plus) platform integrations or Shopify Liquid. Experience across multi brand, multi market, or multi entity technology estates. Familiarity with the GCP ecosystem (BigQuery, Cloud Functions, and related services). Experience integrating with ERP systems (e.g. Microsoft Dynamics 365 Business Central) or fulfilment/WMS platforms. Experience with modern frontend frameworks such as React. Familiarity with SCSS, Tailwind, or modern CSS frameworks. Experience with internal tooling platforms such as Retool. Exposure to e commerce platforms and retail technology ecosystems. What's on offer? Annual bonus scheme Bi Annual Dress Allowance 25 days of annual leave (plus bank holidays) Extra day off for your birthday Flexible working hours around core hours of 10-4 Early Finish Fridays Cycle to work scheme 40% staff discount across Club L and Lavish Alice products Healthcare Cashplan Free onsite gym Enhanced pension contribution Enhanced maternity and sick pay Free snacks, drinks & treats Social events
26/07/2026
Full time
About Us Club L London is the next-generation online fashion retailer for the forward-thinking woman. Conceptualised and crafted in-house and abroad, we specialise in accessible luxury and designs of unrivalled quality that flatter all figures. From prom to occasion, maternity, bridal, and beyond, we deliver an elevated shopping experience that connects our global community of trend-setting consumers, influencers, and content creators with fresh collections dropping weekly. Our engineering team builds and runs the platforms behind Club L London and Lavish Alice across multiple markets - and we're scaling fast. The Role We are looking for a Software Development Manager to lead and grow our engineering team. This is a hands on leadership role that blends people management, delivery ownership, and technical direction. You'll be responsible for the developers who build and maintain our e commerce platforms, integrations, and data systems - setting the standards, shaping the roadmap, and making sure we ship reliable, scalable software at pace. You'll manage a team of full stack and backend engineers, own delivery end to end, and remain close enough to the code to make sound architectural calls and unblock your team. You'll report to the CTO and work across a broad, integration dense stack spanning Shopify Plus, Node.js services, serverless, and a modern GCP/BigQuery data platform. Key Responsibilities Team Leadership & People Management Lead, mentor, and line manage a team of full stack and backend engineers, owning their growth, performance, and development. Run regular 1:1s, set clear objectives, and give consistent, actionable feedback. Drive hiring - define roles, interview, and build a high performing, collaborative engineering culture. Manage team capacity, balance workloads, and protect focus time so the team can deliver sustainably. Delivery & Programme Management Own delivery of the engineering roadmap, translating business priorities into a clear, achievable plan. Run agile ceremonies (planning, stand ups, retrospectives) and keep work moving through the pipeline. Balance BAU, incident response, and project work, managing dependencies, risk, and trade offs. Report progress, blockers, and delivery forecasts clearly to the CTO and wider stakeholders. Technical Leadership & Architecture Guide architecture across e commerce platforms, internal tools, and data systems, ensuring reliability, performance, scalability, and security. Set direction for backend services and APIs (Node.js, REST/GraphQL), serverless functions, and microservices (AWS Lambda or equivalent). Oversee data infrastructure and pipelines across platforms such as BigQuery, MongoDB, or PostgreSQL, ensuring efficient data models and performant queries. Establish and enforce engineering standards - code review culture, testing, CI/CD, and modern development workflows. Platform Integrations & Automation Own the integration landscape connecting e commerce platforms, ERP, marketing, fulfilment, and third party services. Identify opportunities for automation and system improvement across the business, and turn them into delivered outcomes. Ensure integrations are robust, well tested, observable, and resilient as the business scales. Manage technical relationships with key vendors and platform partners. Cross Functional Collaboration & Stakeholder Management Work closely with e commerce, design, content, operations, data, and finance teams to deliver reliable systems and features. Translate business needs into technical delivery, and communicate technical concepts clearly to non technical stakeholders. Champion engineering best practice across the wider tech function. Hands On Contribution Stay technically credible - contribute to code, architecture, code reviews, and complex problem solving where it adds the most value. Lead by example, setting the bar for quality, pragmatism, and delivery. More About You 5+ years in software engineering, with 2+ years leading or line managing engineers, ideally within retail, e commerce, or high traffic digital environments. Proven people leadership - mentoring, performance management, and hiring high performing teams. Strong hands on background with Node.js or similar backend technologies in production. Experience designing and building APIs (REST or GraphQL) and integrating with third party systems. Hands on experience with serverless architectures (AWS Lambda or equivalent). Experience with databases and data platforms such as MongoDB, BigQuery, PostgreSQL, or similar. Solid grounding in CI/CD pipelines, version control (Git), and modern development workflows. A track record of delivering software in an agile environment - planning, prioritising, and shipping reliably. Comfortable balancing hands on technical work with the demands of leading a team. Nice To Have Experience with Shopify (Plus) platform integrations or Shopify Liquid. Experience across multi brand, multi market, or multi entity technology estates. Familiarity with the GCP ecosystem (BigQuery, Cloud Functions, and related services). Experience integrating with ERP systems (e.g. Microsoft Dynamics 365 Business Central) or fulfilment/WMS platforms. Experience with modern frontend frameworks such as React. Familiarity with SCSS, Tailwind, or modern CSS frameworks. Experience with internal tooling platforms such as Retool. Exposure to e commerce platforms and retail technology ecosystems. What's on offer? Annual bonus scheme Bi Annual Dress Allowance 25 days of annual leave (plus bank holidays) Extra day off for your birthday Flexible working hours around core hours of 10-4 Early Finish Fridays Cycle to work scheme 40% staff discount across Club L and Lavish Alice products Healthcare Cashplan Free onsite gym Enhanced pension contribution Enhanced maternity and sick pay Free snacks, drinks & treats Social events
SRE, London
Omaze
Summary People at Apple don't just build products - they craft the kind of experience that have revolutionized entire industries. The diverse collection of our people and their ideas inspire innovation in everything we do. Imagine what you could do here! Join Apple, and help us leave the world better than we found it. The Apple Services Engineering (ASE) team builds and provides systems and infrastructure that fuel Apple's services (such as iCloud, iTunes, Siri, and Maps). We are the foundation on which Apple's software developers build the products that our customers love. We are looking for passionate and talented Site Reliability Engineers to continue our focus in providing our customers the highest quality Apple Services experience. Our services have to scale globally, stay highly available, and "just work." If you love designing, engineering and running systems and infrastructure that will help millions of customers, then this is the place for you! Description FoundationDB infrastructure is BIG. Operating at our scale, across multiple geographically dispersed data centers and servicing hundreds of millions of users presents unique challenges. As an SRE at Apple, you'll need to solve these problems using data, teamwork, and your own expertise. SREs at Apple own the full infrastructure stack; from device driver performance debugging to content delivery network traffic management - our responsibilities are both broad and deep. FoundationDB runs its systems on Linux. We run a mix of open source, vendor licensed, and internally developed tools to perform functions such as system configuration management, provisioning, software deployment, logging, and monitoring. You'll learn these tools and have opportunities to improve them. Our team is collaborative; we work closely with the development teams we support to deliver the best results for Apple. We think critically and strive to balance the best solution with the need to get things done for each engineering challenge we face. Good ideas are heard and results are rewarded. FoundationDB SRE is a small team with huge scale. We serve as the database for much of CloudKit's use cases, including Mail, Contacts, and Keychain. We serve hundreds of millions of customers every day and are a fundamental piece of the Apple device experience. Responsibilities Our team is responsible for the provisioning, managing, and monitoring of FoundationDB in production across multiple regions and control planes (bare-metal, AWS, and Kubernetes). We develop much of our own automation in Java and Go, including our open source Kubernetes Operator (). We work closely with our dev partners to develop a robust and scalable database, often engaging in projects as a single team. Minimum Qualifications Strong sense of ownership and integrity demonstrated through clear communication and collaboration Experience in managing and scaling distributed systems in a public, private, or hybrid cloud environment The ability to design, author, and release code in languages like (but not limited to) Go, Java or Python Acute drive to automate manual operations and to improve them through repeated iteration Understanding of the Linux Operating System, standard networking protocols, and components Hands on experience managing large numbers of diverse systems with configuration management or software delivery platforms (such as Puppet, Chef, Ansible, and Spinnaker) Experience with deploying, supporting and monitoring new and existing services, platforms, and application stacks Excellent troubleshooting and problem solving skills Experience with scale testing, disaster recovery, and capacity planning Familiarity with microservices architecture and container orchestration with Kubernetes Preferred Qualifications Hands on experience managing large numbers of diverse systems with configuration management or software delivery platforms (such as Puppet, Chef, Ansible, and Spinnaker) Experience with deploying, supporting and monitoring new and existing services, platforms, and application stacks Excellent troubleshooting and problem solving skills Experience with scale testing, disaster recovery, and capacity planning Familiarity with microservices architecture and container orchestration with Kubernetes At Apple, we're not all the same. And that's our greatest strength. We draw on the differences in who we are, what we've experienced and how we think. Because to create products that serve everyone, we believe in including everyone. Therefore, we are committed to treating all applicants fairly and equally. As a registered Disability Confident employer, we will work with applicants to make any reasonable accommodations. Apple will consider for employment all qualified applicants with criminal backgrounds in a manner consistent with applicable law. Learn more At Apple, we believe accessibility is a fundamental human right. You'll find that idea reflected in everything here - in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong. Learn about accessibility in Apple's workplace Role Number:
26/07/2026
Full time
Summary People at Apple don't just build products - they craft the kind of experience that have revolutionized entire industries. The diverse collection of our people and their ideas inspire innovation in everything we do. Imagine what you could do here! Join Apple, and help us leave the world better than we found it. The Apple Services Engineering (ASE) team builds and provides systems and infrastructure that fuel Apple's services (such as iCloud, iTunes, Siri, and Maps). We are the foundation on which Apple's software developers build the products that our customers love. We are looking for passionate and talented Site Reliability Engineers to continue our focus in providing our customers the highest quality Apple Services experience. Our services have to scale globally, stay highly available, and "just work." If you love designing, engineering and running systems and infrastructure that will help millions of customers, then this is the place for you! Description FoundationDB infrastructure is BIG. Operating at our scale, across multiple geographically dispersed data centers and servicing hundreds of millions of users presents unique challenges. As an SRE at Apple, you'll need to solve these problems using data, teamwork, and your own expertise. SREs at Apple own the full infrastructure stack; from device driver performance debugging to content delivery network traffic management - our responsibilities are both broad and deep. FoundationDB runs its systems on Linux. We run a mix of open source, vendor licensed, and internally developed tools to perform functions such as system configuration management, provisioning, software deployment, logging, and monitoring. You'll learn these tools and have opportunities to improve them. Our team is collaborative; we work closely with the development teams we support to deliver the best results for Apple. We think critically and strive to balance the best solution with the need to get things done for each engineering challenge we face. Good ideas are heard and results are rewarded. FoundationDB SRE is a small team with huge scale. We serve as the database for much of CloudKit's use cases, including Mail, Contacts, and Keychain. We serve hundreds of millions of customers every day and are a fundamental piece of the Apple device experience. Responsibilities Our team is responsible for the provisioning, managing, and monitoring of FoundationDB in production across multiple regions and control planes (bare-metal, AWS, and Kubernetes). We develop much of our own automation in Java and Go, including our open source Kubernetes Operator (). We work closely with our dev partners to develop a robust and scalable database, often engaging in projects as a single team. Minimum Qualifications Strong sense of ownership and integrity demonstrated through clear communication and collaboration Experience in managing and scaling distributed systems in a public, private, or hybrid cloud environment The ability to design, author, and release code in languages like (but not limited to) Go, Java or Python Acute drive to automate manual operations and to improve them through repeated iteration Understanding of the Linux Operating System, standard networking protocols, and components Hands on experience managing large numbers of diverse systems with configuration management or software delivery platforms (such as Puppet, Chef, Ansible, and Spinnaker) Experience with deploying, supporting and monitoring new and existing services, platforms, and application stacks Excellent troubleshooting and problem solving skills Experience with scale testing, disaster recovery, and capacity planning Familiarity with microservices architecture and container orchestration with Kubernetes Preferred Qualifications Hands on experience managing large numbers of diverse systems with configuration management or software delivery platforms (such as Puppet, Chef, Ansible, and Spinnaker) Experience with deploying, supporting and monitoring new and existing services, platforms, and application stacks Excellent troubleshooting and problem solving skills Experience with scale testing, disaster recovery, and capacity planning Familiarity with microservices architecture and container orchestration with Kubernetes At Apple, we're not all the same. And that's our greatest strength. We draw on the differences in who we are, what we've experienced and how we think. Because to create products that serve everyone, we believe in including everyone. Therefore, we are committed to treating all applicants fairly and equally. As a registered Disability Confident employer, we will work with applicants to make any reasonable accommodations. Apple will consider for employment all qualified applicants with criminal backgrounds in a manner consistent with applicable law. Learn more At Apple, we believe accessibility is a fundamental human right. You'll find that idea reflected in everything here - in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong. Learn about accessibility in Apple's workplace Role Number:
Hippo Digital Limited
Senior Test Engineer
Hippo Digital Limited Leeds, Yorkshire
About The Role Hippo is a rapidly growing digital consultancy passionate about building and delivering transformative digital solutions. We are recruiting for a Senior Test Engineer to join our Hippo Herd. We work at the intersection of strategy, design, and technology to solve complex problems and create lasting value. Our collaborative and agile culture empowers our teams to make a genuine impact. As a Senior Test Engineer, you will play an important role in making Hippo the best Consultancy out there. You will work in multi-disciplinary teams that build, support, and maintain user-centred digital solutions. You will be responsible for analysing features, designing test parameters, creating customised quality checks, and writing final procedures. You will act as a Senior Consultant to deliver testing services to our clients. Senior Test Engineers at Hippo are experienced practitioners who lead technical deliverables, maintain client relationships, and are passionate about developing and upskilling others. Our solutions empower our customers to build and support secure, scalable, and well-engineered systems beyond traditional boundaries. We leverage deep data insights and continuous innovation to deliver awesome platforms that allow our customers to understand and get the most from their data and digital services. The Senior Test Engineer will be a key player and implementer in this mission. Please note we are looking for candidates who are looking for growth at the Senior level; therefore, the advertised salary band represents the entry point for this level, allowing for clear progression within the role, which we are very keen to support. Your Role in a Nutshell Meet with the product design team to determine testing parameters and write comprehensive test plans. Create test cases and conduct quality assurance and design performance tests using new procedures. Work collaboratively with colleagues to explore, design, and deliver solutions to client problems. Design and manage software development and deployment pipelines, resolving issues and potential bottlenecks. Present prototypes, solutions, and progress to internal and external stakeholders in a clear, concise manner. Build great relationships with your team and stakeholders, ensuring challenges are overcome. Lead in the recruitment, on-boarding, and line management of other engineers. Support other consultants in their professional development. Promote Hippo's Engineering Herd externally through blogs, workshops, or conferences. Skills and experience that you need Essential Experience Extensive experience in functional and non-functional automation testing. Extensive experience across the full testing lifecycle with a proven track record as a QA/Test Engineer. Skilled in testing web and mobile applications, including accessibility, cross-browser, and usability testing. Expert knowledge of testing tools such as Selenium, Specflow, Cucumber, Gherkin, JMeter, and K6. Experience with Containerisation tools like Docker or Kubernetes. Expertise in Microservices design and architecture. Solid experience with RESTful API Services. Mastery of Agile methods and techniques such as Scrum, Kanban, and TDD. Skilled in designing and building frameworks to test native and web applications across various scales. Experience with Infrastructure as Code (CloudFormation, Terraform). Expertise in Public Cloud (AWS or Azure) and DevOps. Proficient in CI/CD tools such as Jenkins, AWS CodeBuild, or Azure DevOps. Excellent verbal/written communication skills for technical and non-technical audiences. Desirable Experience Previous experience of working on GOV/NHS projects. Exposure to Python is highly desirable. What makes us great Alongside working with some of the most talented people out there and on-top of a competitive salary, which we're transparent about from the outset, you can also expect a range of benefits: Contributory Pension Scheme (Hippo 6% & Employee 2%) 25 Days Holiday plus UK Public Holidays Perkbox access for a wide range of discounts Critical illness cover Life assurance and death in service cover Volunteer days Cycle-to-work scheme for the avid cyclists Salary sacrifice electric vehicles scheme Season ticket loans Financial and general wellbeing sessions Flexible benefits scheme with options of: Private health cover Private dental cover Additional company pension contributions Additional holidays (up to an extra 2 days) Wellbeing contribution Charity contributions Tree planting Diversity, Inclusion and Belonging at Hippo At Hippo, we're dedicated to creating a diverse, equitable and inclusive workplace that works for everyone. We understand that having a diverse team unlocks our capacity for innovation, creativity and problem solving. Only by building a community of diverse perspectives, cultures and socio-economic backgrounds can we create an environment where all can contribute and thrive. We actively encourage applications from underrepresented groups including women, ethnic minorities, LGBTQ+, neurodivergent and people with disabilities. We are committed to providing an inclusive and accessible recruitment process that reflects our workplace culture. We are a registered Disability Confident Employer, Mindful Employer, Endometriosis Friendly Employer and a member of the Armed Forces Covenant. Hippo continually strives to remove barriers, provide accommodations and offer reasonable adjustments to ensure equity throughout our practices. Hi, we're Hippo At Hippo, we design with empathy and build for impact. We do this by combining data-informed evidence, human-centred design and software engineering. We're a digital services partner who is genuinely invested in helping our clients thrive as modern organisations. Our delivery methodology is truly agile, from concept to reality, supporting innovation and continuous improvement to achieve your desired outcomes. We firmly believe that technology should serve humanity, not the other way around. We take a human-centred approach to everything we do because we understand that complex problems require a service design approach. This means understanding how users behave and ensuring our solutions work for them in the real world. Our combination of data, design, and engineering delivers bespoke digital services that make a positive and meaningful impact on organisations and society. Have a look through our Website at Our Story & Purpose, Our Culture & Values, Case Studies and Our Solutions & Services to find out some more! Hippo Locations We are headquartered in Leeds and have offices across the UK in Glasgow, Manchester, Birmingham, London and Bristol. We're on the lookout for top talent nationwide but you need to be located within reasonable travelling distance from one of our offices. Given the dynamic nature of a consulting business, you may be required to work on-site at a Hippo office or at an in/out of town client location for a number of days per week (client dependent) and therefore candidates will need to be open/flexible to travel and working on one of those sites at least 2 days per week. We offer a generous relocation support package of up to £8,000 (please ask for terms and conditions) to help make your move a smooth one.
25/07/2026
Full time
About The Role Hippo is a rapidly growing digital consultancy passionate about building and delivering transformative digital solutions. We are recruiting for a Senior Test Engineer to join our Hippo Herd. We work at the intersection of strategy, design, and technology to solve complex problems and create lasting value. Our collaborative and agile culture empowers our teams to make a genuine impact. As a Senior Test Engineer, you will play an important role in making Hippo the best Consultancy out there. You will work in multi-disciplinary teams that build, support, and maintain user-centred digital solutions. You will be responsible for analysing features, designing test parameters, creating customised quality checks, and writing final procedures. You will act as a Senior Consultant to deliver testing services to our clients. Senior Test Engineers at Hippo are experienced practitioners who lead technical deliverables, maintain client relationships, and are passionate about developing and upskilling others. Our solutions empower our customers to build and support secure, scalable, and well-engineered systems beyond traditional boundaries. We leverage deep data insights and continuous innovation to deliver awesome platforms that allow our customers to understand and get the most from their data and digital services. The Senior Test Engineer will be a key player and implementer in this mission. Please note we are looking for candidates who are looking for growth at the Senior level; therefore, the advertised salary band represents the entry point for this level, allowing for clear progression within the role, which we are very keen to support. Your Role in a Nutshell Meet with the product design team to determine testing parameters and write comprehensive test plans. Create test cases and conduct quality assurance and design performance tests using new procedures. Work collaboratively with colleagues to explore, design, and deliver solutions to client problems. Design and manage software development and deployment pipelines, resolving issues and potential bottlenecks. Present prototypes, solutions, and progress to internal and external stakeholders in a clear, concise manner. Build great relationships with your team and stakeholders, ensuring challenges are overcome. Lead in the recruitment, on-boarding, and line management of other engineers. Support other consultants in their professional development. Promote Hippo's Engineering Herd externally through blogs, workshops, or conferences. Skills and experience that you need Essential Experience Extensive experience in functional and non-functional automation testing. Extensive experience across the full testing lifecycle with a proven track record as a QA/Test Engineer. Skilled in testing web and mobile applications, including accessibility, cross-browser, and usability testing. Expert knowledge of testing tools such as Selenium, Specflow, Cucumber, Gherkin, JMeter, and K6. Experience with Containerisation tools like Docker or Kubernetes. Expertise in Microservices design and architecture. Solid experience with RESTful API Services. Mastery of Agile methods and techniques such as Scrum, Kanban, and TDD. Skilled in designing and building frameworks to test native and web applications across various scales. Experience with Infrastructure as Code (CloudFormation, Terraform). Expertise in Public Cloud (AWS or Azure) and DevOps. Proficient in CI/CD tools such as Jenkins, AWS CodeBuild, or Azure DevOps. Excellent verbal/written communication skills for technical and non-technical audiences. Desirable Experience Previous experience of working on GOV/NHS projects. Exposure to Python is highly desirable. What makes us great Alongside working with some of the most talented people out there and on-top of a competitive salary, which we're transparent about from the outset, you can also expect a range of benefits: Contributory Pension Scheme (Hippo 6% & Employee 2%) 25 Days Holiday plus UK Public Holidays Perkbox access for a wide range of discounts Critical illness cover Life assurance and death in service cover Volunteer days Cycle-to-work scheme for the avid cyclists Salary sacrifice electric vehicles scheme Season ticket loans Financial and general wellbeing sessions Flexible benefits scheme with options of: Private health cover Private dental cover Additional company pension contributions Additional holidays (up to an extra 2 days) Wellbeing contribution Charity contributions Tree planting Diversity, Inclusion and Belonging at Hippo At Hippo, we're dedicated to creating a diverse, equitable and inclusive workplace that works for everyone. We understand that having a diverse team unlocks our capacity for innovation, creativity and problem solving. Only by building a community of diverse perspectives, cultures and socio-economic backgrounds can we create an environment where all can contribute and thrive. We actively encourage applications from underrepresented groups including women, ethnic minorities, LGBTQ+, neurodivergent and people with disabilities. We are committed to providing an inclusive and accessible recruitment process that reflects our workplace culture. We are a registered Disability Confident Employer, Mindful Employer, Endometriosis Friendly Employer and a member of the Armed Forces Covenant. Hippo continually strives to remove barriers, provide accommodations and offer reasonable adjustments to ensure equity throughout our practices. Hi, we're Hippo At Hippo, we design with empathy and build for impact. We do this by combining data-informed evidence, human-centred design and software engineering. We're a digital services partner who is genuinely invested in helping our clients thrive as modern organisations. Our delivery methodology is truly agile, from concept to reality, supporting innovation and continuous improvement to achieve your desired outcomes. We firmly believe that technology should serve humanity, not the other way around. We take a human-centred approach to everything we do because we understand that complex problems require a service design approach. This means understanding how users behave and ensuring our solutions work for them in the real world. Our combination of data, design, and engineering delivers bespoke digital services that make a positive and meaningful impact on organisations and society. Have a look through our Website at Our Story & Purpose, Our Culture & Values, Case Studies and Our Solutions & Services to find out some more! Hippo Locations We are headquartered in Leeds and have offices across the UK in Glasgow, Manchester, Birmingham, London and Bristol. We're on the lookout for top talent nationwide but you need to be located within reasonable travelling distance from one of our offices. Given the dynamic nature of a consulting business, you may be required to work on-site at a Hippo office or at an in/out of town client location for a number of days per week (client dependent) and therefore candidates will need to be open/flexible to travel and working on one of those sites at least 2 days per week. We offer a generous relocation support package of up to £8,000 (please ask for terms and conditions) to help make your move a smooth one.
Cloud Platform Engineer (Senior / Lead)
SimplyBiz PLC
Cloud Platform Engineer (Senior / Lead) Department: Technology Employment Type: Permanent - Full Time Location: London Reporting To: Description About the role We are looking for a hands on Senior / Lead Cloud Platform Engineer to shape and run the multi cloud foundation that our engineering, data and product teams build on everyday. This is a senior individual contributor / tech lead role at the heart of a modern hybrid cloud strategy. Our goal is to make cloud infrastructure effectively invisible and commoditised for internal teams: engineers should ship through self service golden paths and a well designed internal developer platform, without needing to understand the underlying cloud plumbing. You will treat platform capabilities as products, with clear APIs, paved roads, strong defaults and great developer experience. A defining part of this role is preparing our infrastructure for a future (or present) where some "engineers" are AI agents. You will design platforms, guardrails and interfaces that let both human engineers and autonomous AI agents provision, operate and optimise infrastructure safely - and you will use AI agents heavily yourself to automate and optimise day to day platform work. Security is central, not an afterthought. You will help drive a zero trust approach across identity, network, workloads and data, ensuring the platform is secure by default for humans and machine/agent identities alike. You will also work closely with our data function, helping design and optimise the data pipelines, data stores and large scale analytics infrastructure (for example BigQuery and similar warehouses) that the business depends on. What you'll do Design, build and operate our multi cloud and hybrid cloud platform across at least two of the top three providers (AWS, Azure and/or Google Cloud), plus on prem/hybrid connectivity where needed. Build and own an internal developer platform and self service "golden paths" that make cloud infrastructure feel invisible and commoditised for engineering, data and product teams; and their AI agents. Deliver everything as code: infrastructure as code, GitOps, reusable modules, CI/CD pipelines and policy as code guardrails. Leverage AI agents extensively to automate and optimise platform work-provisioning, cost and performance optimisation, incident response, remediation and documentation. Prepare the infrastructure for AI agents as first class "engineers": safe machine identities, scoped permissions, sandboxes, approval workflows and audit trails so agents can provision and operate infrastructure within tight guardrails. Embed a zero trust security model across identity, network, workloads and data for both human and machine/agent identities; secure by default, least privilege, secrets management and continuous compliance. Apply SRE practices-SLOs/SLIs, observability, capacity planning, resilience and blameless incident management-to keep the platform reliable and cost efficient. Partner with data engineering to design and optimise data pipelines, data stores and large scale analytics infrastructure such as BigQuery, including query, cost and performance tuning. Mentor engineers, set technical direction and champion strong platform and security engineering standards across the organisation. What you'll need to succeed Essential requirements: Extensive hands on experience designing, building and operating production cloud infrastructure at senior or lead level. Multi cloud experience across at least two of the top three providers (AWS, Microsoft Azure and Google Cloud), including a recognised professional level cloud certification for each of those two providers (for example AWS Solutions Architect / DevOps Engineer Professional, Azure Solutions Architect / DevOps Engineer Expert, or Google Cloud Professional Cloud Architect / DevOps Engineer). Strong background in modern hybrid cloud architecture and connecting cloud with on prem/edge environments. Deep infrastructure as code and automation skills (e.g. Terraform/OpenTofu, Pulumi, Ansible), GitOps and CI/CD, plus containers and orchestration (Docker, Kubernetes). Proven experience building internal developer platforms, self service golden paths and platform as a product to abstract away cloud complexity for engineering teams. Practical experience using AI agents / LLM based tooling to automate and optimise infrastructure work, and interest in designing infrastructure that AI agents can operate safely. Strong security engineering mindset with hands on zero trust experience across identity, network, workloads and data-including secrets management, least privilege IAM and machine/workload identity. Solid programming/scripting ability (e.g. Python, Go) and strong observability, reliability and cost optimisation practices. Desirable requirements: Experience working as a Site Reliability Engineer (SRE) with SLOs/SLIs, error budgets and incident management. A third top tier cloud certification, or specialist security/Kubernetes certifications (e.g. CKA/CKS). Significant data engineering experience: designing and operating data pipelines and data stores, and optimising databases and large scale data infrastructure such as BigQuery (including query, cost and performance tuning). Experience preparing environments for autonomous or agentic workloads-sandboxes, scoped machine identities, approval workflows and audit trails. Experience in a regulated or fintech environment. Your approach to work: Pragmatic and hands on, with a strong bias for automation and eliminating toil. Product mindset- you treat internal engineers (human and AI) as your customers and obsess over their experience. Security and reliability first, collaborative, and comfortable leading and mentoring. Important to know Location: We have multiple offices across the UK. We have a new office in London which is becoming more central to where we collaborate in person. We have a flexible working policy with a few days per week in the office. Right to Work: Applicants must already hold a legal right to work in the UK without time restrictions and without the need for future sponsorship. We are unable to provide Skilled Worker visa sponsorship.
25/07/2026
Full time
Cloud Platform Engineer (Senior / Lead) Department: Technology Employment Type: Permanent - Full Time Location: London Reporting To: Description About the role We are looking for a hands on Senior / Lead Cloud Platform Engineer to shape and run the multi cloud foundation that our engineering, data and product teams build on everyday. This is a senior individual contributor / tech lead role at the heart of a modern hybrid cloud strategy. Our goal is to make cloud infrastructure effectively invisible and commoditised for internal teams: engineers should ship through self service golden paths and a well designed internal developer platform, without needing to understand the underlying cloud plumbing. You will treat platform capabilities as products, with clear APIs, paved roads, strong defaults and great developer experience. A defining part of this role is preparing our infrastructure for a future (or present) where some "engineers" are AI agents. You will design platforms, guardrails and interfaces that let both human engineers and autonomous AI agents provision, operate and optimise infrastructure safely - and you will use AI agents heavily yourself to automate and optimise day to day platform work. Security is central, not an afterthought. You will help drive a zero trust approach across identity, network, workloads and data, ensuring the platform is secure by default for humans and machine/agent identities alike. You will also work closely with our data function, helping design and optimise the data pipelines, data stores and large scale analytics infrastructure (for example BigQuery and similar warehouses) that the business depends on. What you'll do Design, build and operate our multi cloud and hybrid cloud platform across at least two of the top three providers (AWS, Azure and/or Google Cloud), plus on prem/hybrid connectivity where needed. Build and own an internal developer platform and self service "golden paths" that make cloud infrastructure feel invisible and commoditised for engineering, data and product teams; and their AI agents. Deliver everything as code: infrastructure as code, GitOps, reusable modules, CI/CD pipelines and policy as code guardrails. Leverage AI agents extensively to automate and optimise platform work-provisioning, cost and performance optimisation, incident response, remediation and documentation. Prepare the infrastructure for AI agents as first class "engineers": safe machine identities, scoped permissions, sandboxes, approval workflows and audit trails so agents can provision and operate infrastructure within tight guardrails. Embed a zero trust security model across identity, network, workloads and data for both human and machine/agent identities; secure by default, least privilege, secrets management and continuous compliance. Apply SRE practices-SLOs/SLIs, observability, capacity planning, resilience and blameless incident management-to keep the platform reliable and cost efficient. Partner with data engineering to design and optimise data pipelines, data stores and large scale analytics infrastructure such as BigQuery, including query, cost and performance tuning. Mentor engineers, set technical direction and champion strong platform and security engineering standards across the organisation. What you'll need to succeed Essential requirements: Extensive hands on experience designing, building and operating production cloud infrastructure at senior or lead level. Multi cloud experience across at least two of the top three providers (AWS, Microsoft Azure and Google Cloud), including a recognised professional level cloud certification for each of those two providers (for example AWS Solutions Architect / DevOps Engineer Professional, Azure Solutions Architect / DevOps Engineer Expert, or Google Cloud Professional Cloud Architect / DevOps Engineer). Strong background in modern hybrid cloud architecture and connecting cloud with on prem/edge environments. Deep infrastructure as code and automation skills (e.g. Terraform/OpenTofu, Pulumi, Ansible), GitOps and CI/CD, plus containers and orchestration (Docker, Kubernetes). Proven experience building internal developer platforms, self service golden paths and platform as a product to abstract away cloud complexity for engineering teams. Practical experience using AI agents / LLM based tooling to automate and optimise infrastructure work, and interest in designing infrastructure that AI agents can operate safely. Strong security engineering mindset with hands on zero trust experience across identity, network, workloads and data-including secrets management, least privilege IAM and machine/workload identity. Solid programming/scripting ability (e.g. Python, Go) and strong observability, reliability and cost optimisation practices. Desirable requirements: Experience working as a Site Reliability Engineer (SRE) with SLOs/SLIs, error budgets and incident management. A third top tier cloud certification, or specialist security/Kubernetes certifications (e.g. CKA/CKS). Significant data engineering experience: designing and operating data pipelines and data stores, and optimising databases and large scale data infrastructure such as BigQuery (including query, cost and performance tuning). Experience preparing environments for autonomous or agentic workloads-sandboxes, scoped machine identities, approval workflows and audit trails. Experience in a regulated or fintech environment. Your approach to work: Pragmatic and hands on, with a strong bias for automation and eliminating toil. Product mindset- you treat internal engineers (human and AI) as your customers and obsess over their experience. Security and reliability first, collaborative, and comfortable leading and mentoring. Important to know Location: We have multiple offices across the UK. We have a new office in London which is becoming more central to where we collaborate in person. We have a flexible working policy with a few days per week in the office. Right to Work: Applicants must already hold a legal right to work in the UK without time restrictions and without the need for future sponsorship. We are unable to provide Skilled Worker visa sponsorship.
Amazon
European Backbone Network Architect
Amazon
Backbone Fiber - Business & Vendor Developer , Global Connectivity & Infrastructure Development Job ID: Amazon Development Centre (London) Limited This position can be based in Dublin or London. Today, Amazon Web Services provides a highly reliable, scalable, low-cost infrastructure platform in the cloud that powers hundreds of thousands of businesses in 190 countries around the world. The AWS Cloud infrastructure is built around Regions and Availability Zones (AZs). AWS Regions provide multiple, physically separated and isolated Availability Zones which are connected with low latency, high throughput, and highly redundant networking. These Availability Zones offer AWS customers an easier and more effective way to design and operate applications and databases, making them more highly available, fault tolerant, and scalable than traditional single datacenter infrastructures or multi-datacenter infrastructures. How would you like to come be part of the team that builds out that low latency, high throughput, and highly redundant network? AWS is seeking an exceptional Backbone Network Developer to help architect and scale one of the world's most sophisticated networks. As a key member of our Backbone Network Development organization, you will drive innovation in network scaling while maintaining AWS's renowned operational excellence. This role offers the unique opportunity to shape the future of global network infrastructure while solving complex challenges at unprecedented scale. You will leverage your technical expertise and commercial acumen to develop creative solutions that support Amazon's continued worldwide growth and expansion. Join us in building the future of cloud infrastructure at AWS, where your contributions will have global impact. This position sits within our Backbone Network Development organization, a team critical to Amazon's global infrastructure and connectivity strategy. AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other words, we're the people who keep the cloud running. We support all AWS data centers and all of the servers, storage, networking, power, and cooling equipment that ensure our customers have continual access to the innovation they rely on. We work on the most challenging problems, with thousands of variables impacting the supply chain - and we're looking for talented people who want to help. You'll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers, and other vital roles. You'll collaborate with people across AWS to help us deliver the highest standards for safety and security while providing seemingly infinite capacity at the lowest possible cost for our customers. And you'll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion. Key job responsibilities Network Strategy & Development - Drive AWS's European backbone network infrastructure strategy and execution - Design and implement large-scale network architecture supporting Amazon's global operations - Develop and execute strategic network expansion plans with regular milestone tracking - Optimize network topology for maximum efficiency and cost-effectiveness - Monitor and analyze telecommunications market trends, emerging technologies, and industry developments Vendor & Partnership Management - Identify and evaluate strategic suppliers for wavelengths, dark fiber, and submarine cable capacity - Lead technical vendor strategy across AWS's European network topology - Develop and maintain key commercial and technical partnerships within Europe - Manage supplier onboarding, performance metrics, and service quality standards - Coordinate with vendors and internal teams to deliver new backbone capacity Financial & Project Management - Collaborate with finance partners on budget planning and allocation - Manage risk assessment, escalations, and technical constraint resolution - Balance business requirements with technical limitations and cost considerations About the team Diverse Experiences AWS values diverse experiences. Even if you do not meet all of the qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn't followed a traditional path, or includes alternative experiences, don't let it stop you from applying. Why AWS? Amazon Web Services (AWS) is the world's most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating - that's why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses. Inclusive Team Culture Here at AWS, it's in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon conferences, inspire us to never stop embracing our uniqueness. Mentorship & Career Growth We're continuously raising our performance bar as we strive to become Earth's Best Employer. That's why you'll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional. Work/Life Balance We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there's nothing we can't achieve in the cloud. Basic Qualifications - Bachelor's degree in Information Technology, Computer Science, or a related field, or experience related IT - Experience related to spectrum management, securing licenses for the provision of telecommunications services and ensuring regulatory compliance at an international level - Experience managing procurement teams with direct experience in supplier/supply chain management - Experience in solving complex business challenges by delivering accurate and timely financial models, analysis, and recommendations that have a proven impact on business (e.g., financial savings, operational improvements, or customer benefits), or experience building and managing financial models for business forecasting and problem solving Preferred Qualifications - Experience working in a network planning, network construction/implementation, or commercial implementation role at an ISP, telecoms operator, or Hyperscaler - Experience in strategic marketing management and market analysis and demonstrated ability to build and execute a strategy with clear goals and objectives to align to business and service objectives, and support portfolio objectives - Experience working across functional teams and senior stakeholders - Experience in Network protocols like DNS/DHCP/TCP, or experience in Linux and Networking protocols and experience that includes strong analytical skills, attention to detail, and effective communication abilities Amazon is an equal opportunities employer. We believe passionately that employing a diverse workforce is central to our success. We make recruiting decisions based on your experience and skills. We value your passion to discover, invent, simplify and build. Protecting your privacy and the security of your data is a longstanding top priority for Amazon. Please consult our Privacy Notice ( ) to know more about how we collect, use and transfer the personal data of our candidates. Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. Posted: June 15, 2026 (Updated 1 day ago) Posted: June 25, 2026 (Updated 3 days ago) Posted: July 13, 2026 (Updated 4 days ago) Posted: April 29, 2026 (Updated 4 days ago) Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
25/07/2026
Full time
Backbone Fiber - Business & Vendor Developer , Global Connectivity & Infrastructure Development Job ID: Amazon Development Centre (London) Limited This position can be based in Dublin or London. Today, Amazon Web Services provides a highly reliable, scalable, low-cost infrastructure platform in the cloud that powers hundreds of thousands of businesses in 190 countries around the world. The AWS Cloud infrastructure is built around Regions and Availability Zones (AZs). AWS Regions provide multiple, physically separated and isolated Availability Zones which are connected with low latency, high throughput, and highly redundant networking. These Availability Zones offer AWS customers an easier and more effective way to design and operate applications and databases, making them more highly available, fault tolerant, and scalable than traditional single datacenter infrastructures or multi-datacenter infrastructures. How would you like to come be part of the team that builds out that low latency, high throughput, and highly redundant network? AWS is seeking an exceptional Backbone Network Developer to help architect and scale one of the world's most sophisticated networks. As a key member of our Backbone Network Development organization, you will drive innovation in network scaling while maintaining AWS's renowned operational excellence. This role offers the unique opportunity to shape the future of global network infrastructure while solving complex challenges at unprecedented scale. You will leverage your technical expertise and commercial acumen to develop creative solutions that support Amazon's continued worldwide growth and expansion. Join us in building the future of cloud infrastructure at AWS, where your contributions will have global impact. This position sits within our Backbone Network Development organization, a team critical to Amazon's global infrastructure and connectivity strategy. AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other words, we're the people who keep the cloud running. We support all AWS data centers and all of the servers, storage, networking, power, and cooling equipment that ensure our customers have continual access to the innovation they rely on. We work on the most challenging problems, with thousands of variables impacting the supply chain - and we're looking for talented people who want to help. You'll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers, and other vital roles. You'll collaborate with people across AWS to help us deliver the highest standards for safety and security while providing seemingly infinite capacity at the lowest possible cost for our customers. And you'll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion. Key job responsibilities Network Strategy & Development - Drive AWS's European backbone network infrastructure strategy and execution - Design and implement large-scale network architecture supporting Amazon's global operations - Develop and execute strategic network expansion plans with regular milestone tracking - Optimize network topology for maximum efficiency and cost-effectiveness - Monitor and analyze telecommunications market trends, emerging technologies, and industry developments Vendor & Partnership Management - Identify and evaluate strategic suppliers for wavelengths, dark fiber, and submarine cable capacity - Lead technical vendor strategy across AWS's European network topology - Develop and maintain key commercial and technical partnerships within Europe - Manage supplier onboarding, performance metrics, and service quality standards - Coordinate with vendors and internal teams to deliver new backbone capacity Financial & Project Management - Collaborate with finance partners on budget planning and allocation - Manage risk assessment, escalations, and technical constraint resolution - Balance business requirements with technical limitations and cost considerations About the team Diverse Experiences AWS values diverse experiences. Even if you do not meet all of the qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn't followed a traditional path, or includes alternative experiences, don't let it stop you from applying. Why AWS? Amazon Web Services (AWS) is the world's most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating - that's why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses. Inclusive Team Culture Here at AWS, it's in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon conferences, inspire us to never stop embracing our uniqueness. Mentorship & Career Growth We're continuously raising our performance bar as we strive to become Earth's Best Employer. That's why you'll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional. Work/Life Balance We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there's nothing we can't achieve in the cloud. Basic Qualifications - Bachelor's degree in Information Technology, Computer Science, or a related field, or experience related IT - Experience related to spectrum management, securing licenses for the provision of telecommunications services and ensuring regulatory compliance at an international level - Experience managing procurement teams with direct experience in supplier/supply chain management - Experience in solving complex business challenges by delivering accurate and timely financial models, analysis, and recommendations that have a proven impact on business (e.g., financial savings, operational improvements, or customer benefits), or experience building and managing financial models for business forecasting and problem solving Preferred Qualifications - Experience working in a network planning, network construction/implementation, or commercial implementation role at an ISP, telecoms operator, or Hyperscaler - Experience in strategic marketing management and market analysis and demonstrated ability to build and execute a strategy with clear goals and objectives to align to business and service objectives, and support portfolio objectives - Experience working across functional teams and senior stakeholders - Experience in Network protocols like DNS/DHCP/TCP, or experience in Linux and Networking protocols and experience that includes strong analytical skills, attention to detail, and effective communication abilities Amazon is an equal opportunities employer. We believe passionately that employing a diverse workforce is central to our success. We make recruiting decisions based on your experience and skills. We value your passion to discover, invent, simplify and build. Protecting your privacy and the security of your data is a longstanding top priority for Amazon. Please consult our Privacy Notice ( ) to know more about how we collect, use and transfer the personal data of our candidates. Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner. Posted: June 15, 2026 (Updated 1 day ago) Posted: June 25, 2026 (Updated 3 days ago) Posted: July 13, 2026 (Updated 4 days ago) Posted: April 29, 2026 (Updated 4 days ago) Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Principal Engineer, Product Software
United States Digital Space LLC
Principal Engineer, Product Software JR-162356 Hybrid London Technology Full time Who are we? the company is the world's digital infrastructure company , shortening the path to connectivity to enable the innovations that enrich our work, life and planet. A place where bold ideas are welcomed, human connection is valued, and everyone has the opportunity to shape their future. Help us challenge assumptions, uncover bias, and remove barriers-because progress starts with fresh ideas. You'll find belonging, purpose, and a team that welcomes you-because when you feel valued, you're empowered to do your best work. Job Summary the company Incubation builds agentic systems that support the software delivery lifecycle end to end through specialized AI agents, human oversight, and enterprise-grade governance. As a Principal Engineer, you will provide technical leadership for the architecture, design, and implementation of critical platform and agent capabilities. This is the most senior individual contributor role within the organization and is responsible for solving complex technical challenges, establishing engineering standards, and influencing the technical direction of the platform. This role may focus on either the Agent Platform & Orchestration domain or the Agent Engineering domain, depending on business needs. Regardless of assignment, you will serve as a technical leader, mentor, and trusted advisor across engineering teams. Architecture and Technical Leadership Design scalable, secure, and reliable architectures supporting agent orchestration, agent-to-agent communication, tool integration, event-driven workflows, and enterprise platform integration. Lead the resolution of complex technical challenges related to reliability, performance, scalability, evaluation frameworks, and operational excellence. Develop and review critical platform and application code, ensuring adherence to software engineering best practices. Provide technical guidance and architectural direction across multiple engineering teams. Influence the technical strategy and long term architecture of the platform. Engineering Excellence and Quality Define engineering standards, evaluation methodologies, guardrails, release criteria, and quality frameworks for agent-based systems. Establish observability, monitoring, and evaluation capabilities that enable measurable and repeatable quality outcomes. Review and approve key technical designs and architectural decisions. Promote engineering rigor through technical reviews, documentation, and standards adoption. Ensure technical quality, maintainability, and operational readiness across platform and agent capabilities. Platform and Product Delivery Partner with product management and engineering leadership to translate business objectives into technical strategy and execution plans. Create reusable frameworks, patterns, and reference implementations that accelerate platform development and adoption. Drive continuous improvement of platform capabilities, reliability, operational efficiency, and developer experience. Balance technical innovation with business priorities, delivery timelines, and operational requirements. Technical Mentorship and Influence Mentor and develop engineers through technical coaching, design reviews, and knowledge sharing. Foster adoption of engineering standards, best practices, and platform patterns across teams. Influence technical decision-making through expertise, collaboration, and leadership without direct authority. Build engineering capability by promoting high standards of technical excellence and craftsmanship. Security, Governance, and Operations Ensure platform designs incorporate security, privacy, compliance, auditability, and responsible AI considerations. Partner with security, governance, and architecture stakeholders to meet enterprise requirements. Drive operational excellence through monitoring, resilience, capacity planning, cost optimization, and incident readiness. Champion secure by design and reliable by design engineering practices across the platform. Required Extensive software engineering experience with a proven record of technical leadership across multiple teams and organizations. Deep experience designing, building, and operating distributed systems, cloud-native platforms, APIs, and event driven architectures. Hands on experience building and delivering LLM powered applications, AI systems, or agent based solutions in production environments. Strong software development expertise in Python, TypeScript, or comparable modern programming languages. Demonstrated experience establishing technical standards, engineering practices, and architectural direction at scale. Proven ability to mentor senior engineers and develop technical talent. Strong written and verbal communication skills with the ability to influence technical and business stakeholders. Preferred Qualifications Experience with agent frameworks, orchestration platforms, and protocols such as MCP, A2A, LangGraph, Bedrock, Vertex AI, OpenAI, Anthropic, or similar technologies. Experience building developer platforms, orchestration engines, SDLC tooling, or engineering productivity solutions. Experience implementing enterprise AI systems within regulated or security-conscious environments. Experience with responsible AI governance, evaluation frameworks, and model quality practices. Contributions to open-source software, technical publications, patents, or industry conference presentations. the company Values At the company, we operate with a growth mindset and embrace diversity of thought and contribution. We are committed to creating an environment where everyone can do their best work, grow their careers, and contribute to our success. We welcome people with diverse experiences and backgrounds who are passionate about solving complex technical challenges and building technology that creates meaningful business impact. the company is committed to ensuring that our employment process is open to all individuals, including those with a disability. If you are a qualified candidate and need assistance or an accommodation, please let us know by completing form. the company is an Equal Employment Opportunity and, in the U.S., an Affir mative Action employer. All qualified applicants will receive consideration for employment without regard to unlawful consideration of race, color, religion, creed, national or ethnic origin, ancestry, place of birth, citizenship, sex, pregnancy / childbirth or related medical conditions, sexual orientation, gender identity or expression, marital or domestic partnership status, age, veteran or military status, physical or mental disability, medical condition, genetic information, political / organizational affiliation, status as a victim or family member of a victim of crime or abuse, or any other status protected by applicable law. We use artificial intelligence in our hiring process. Learn more here. This posting is a new position within our organization.
25/07/2026
Full time
Principal Engineer, Product Software JR-162356 Hybrid London Technology Full time Who are we? the company is the world's digital infrastructure company , shortening the path to connectivity to enable the innovations that enrich our work, life and planet. A place where bold ideas are welcomed, human connection is valued, and everyone has the opportunity to shape their future. Help us challenge assumptions, uncover bias, and remove barriers-because progress starts with fresh ideas. You'll find belonging, purpose, and a team that welcomes you-because when you feel valued, you're empowered to do your best work. Job Summary the company Incubation builds agentic systems that support the software delivery lifecycle end to end through specialized AI agents, human oversight, and enterprise-grade governance. As a Principal Engineer, you will provide technical leadership for the architecture, design, and implementation of critical platform and agent capabilities. This is the most senior individual contributor role within the organization and is responsible for solving complex technical challenges, establishing engineering standards, and influencing the technical direction of the platform. This role may focus on either the Agent Platform & Orchestration domain or the Agent Engineering domain, depending on business needs. Regardless of assignment, you will serve as a technical leader, mentor, and trusted advisor across engineering teams. Architecture and Technical Leadership Design scalable, secure, and reliable architectures supporting agent orchestration, agent-to-agent communication, tool integration, event-driven workflows, and enterprise platform integration. Lead the resolution of complex technical challenges related to reliability, performance, scalability, evaluation frameworks, and operational excellence. Develop and review critical platform and application code, ensuring adherence to software engineering best practices. Provide technical guidance and architectural direction across multiple engineering teams. Influence the technical strategy and long term architecture of the platform. Engineering Excellence and Quality Define engineering standards, evaluation methodologies, guardrails, release criteria, and quality frameworks for agent-based systems. Establish observability, monitoring, and evaluation capabilities that enable measurable and repeatable quality outcomes. Review and approve key technical designs and architectural decisions. Promote engineering rigor through technical reviews, documentation, and standards adoption. Ensure technical quality, maintainability, and operational readiness across platform and agent capabilities. Platform and Product Delivery Partner with product management and engineering leadership to translate business objectives into technical strategy and execution plans. Create reusable frameworks, patterns, and reference implementations that accelerate platform development and adoption. Drive continuous improvement of platform capabilities, reliability, operational efficiency, and developer experience. Balance technical innovation with business priorities, delivery timelines, and operational requirements. Technical Mentorship and Influence Mentor and develop engineers through technical coaching, design reviews, and knowledge sharing. Foster adoption of engineering standards, best practices, and platform patterns across teams. Influence technical decision-making through expertise, collaboration, and leadership without direct authority. Build engineering capability by promoting high standards of technical excellence and craftsmanship. Security, Governance, and Operations Ensure platform designs incorporate security, privacy, compliance, auditability, and responsible AI considerations. Partner with security, governance, and architecture stakeholders to meet enterprise requirements. Drive operational excellence through monitoring, resilience, capacity planning, cost optimization, and incident readiness. Champion secure by design and reliable by design engineering practices across the platform. Required Extensive software engineering experience with a proven record of technical leadership across multiple teams and organizations. Deep experience designing, building, and operating distributed systems, cloud-native platforms, APIs, and event driven architectures. Hands on experience building and delivering LLM powered applications, AI systems, or agent based solutions in production environments. Strong software development expertise in Python, TypeScript, or comparable modern programming languages. Demonstrated experience establishing technical standards, engineering practices, and architectural direction at scale. Proven ability to mentor senior engineers and develop technical talent. Strong written and verbal communication skills with the ability to influence technical and business stakeholders. Preferred Qualifications Experience with agent frameworks, orchestration platforms, and protocols such as MCP, A2A, LangGraph, Bedrock, Vertex AI, OpenAI, Anthropic, or similar technologies. Experience building developer platforms, orchestration engines, SDLC tooling, or engineering productivity solutions. Experience implementing enterprise AI systems within regulated or security-conscious environments. Experience with responsible AI governance, evaluation frameworks, and model quality practices. Contributions to open-source software, technical publications, patents, or industry conference presentations. the company Values At the company, we operate with a growth mindset and embrace diversity of thought and contribution. We are committed to creating an environment where everyone can do their best work, grow their careers, and contribute to our success. We welcome people with diverse experiences and backgrounds who are passionate about solving complex technical challenges and building technology that creates meaningful business impact. the company is committed to ensuring that our employment process is open to all individuals, including those with a disability. If you are a qualified candidate and need assistance or an accommodation, please let us know by completing form. the company is an Equal Employment Opportunity and, in the U.S., an Affir mative Action employer. All qualified applicants will receive consideration for employment without regard to unlawful consideration of race, color, religion, creed, national or ethnic origin, ancestry, place of birth, citizenship, sex, pregnancy / childbirth or related medical conditions, sexual orientation, gender identity or expression, marital or domestic partnership status, age, veteran or military status, physical or mental disability, medical condition, genetic information, political / organizational affiliation, status as a victim or family member of a victim of crime or abuse, or any other status protected by applicable law. We use artificial intelligence in our hiring process. Learn more here. This posting is a new position within our organization.
Amazon
Backbone Fiber - Business & Vendor Developer , Global Connectivity & Infrastructure Development ...
Amazon
Backbone Fiber - Business & Vendor Developer , Global Connectivity & Infrastructure Development Job ID: Amazon Development Centre (London) Limited This position can be based in Dublin or London. Amazon Web Services provides a highly reliable, scalable, low-cost infrastructure platform in the cloud that powers businesses worldwide. The AWS Cloud infrastructure is built around Regions and Availability Zones (AZs). AWS Regions provide multiple, physically separated and isolated Availability Zones connected with low latency, high throughput, and highly redundant networking. AWS is seeking an exceptional Backbone Network Developer to architect and scale one of the world's most sophisticated networks. As a key member of our Backbone Network Development organization, you will drive innovation in network scaling while maintaining AWS's renowned operational excellence. Key Job Responsibilities Drive AWS's European backbone network infrastructure strategy and execution. Design and implement large scale network architecture supporting Amazon's global operations. Develop and execute strategic network expansion plans with regular milestone tracking. Optimize network topology for maximum efficiency and cost effectiveness. Monitor and analyze telecommunications market trends, emerging technologies, and industry developments. Identify and evaluate strategic suppliers for wavelengths, dark fiber, and submarine cable capacity. Lead technical vendor strategy across AWS's European network topology. Develop and maintain key commercial and technical partnerships within Europe. Manage supplier onboarding, performance metrics, and service quality standards. Coordinate with vendors and internal teams to deliver new backbone capacity. Collaborate with finance partners on budget planning and allocation. Manage risk assessment, escalations, and technical constraint resolution. Balance business requirements with technical limitations and cost considerations. Basic Qualifications Bachelor's degree in Information Technology, Computer Science, or a related field, or experience related IT. Experience in spectrum management, securing licensing for telecommunications services, and ensuring regulatory compliance at an international level. Experience managing procurement teams with direct experience in supplier/supply chain management. Experience in solving complex business challenges by delivering accurate and timely financial models, analysis, and recommendations that have a proven impact on business (e.g., financial savings, operational improvements, or customer benefits). Preferred Qualifications Experience in a network planning, construction/implementation, or commercial implementation role at an ISP, telecoms operator, or hyperscaler. Experience in strategic marketing management and market analysis, demonstrated ability to build and execute a strategy with clear goals and objectives aligned to business and service objectives. Experience working across functional teams and senior stakeholders. Experience with network protocols such as DNS, DHCP, TCP, or Linux and networking protocols, with strong analytical skills, attention to detail, and effective communication abilities. Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
25/07/2026
Full time
Backbone Fiber - Business & Vendor Developer , Global Connectivity & Infrastructure Development Job ID: Amazon Development Centre (London) Limited This position can be based in Dublin or London. Amazon Web Services provides a highly reliable, scalable, low-cost infrastructure platform in the cloud that powers businesses worldwide. The AWS Cloud infrastructure is built around Regions and Availability Zones (AZs). AWS Regions provide multiple, physically separated and isolated Availability Zones connected with low latency, high throughput, and highly redundant networking. AWS is seeking an exceptional Backbone Network Developer to architect and scale one of the world's most sophisticated networks. As a key member of our Backbone Network Development organization, you will drive innovation in network scaling while maintaining AWS's renowned operational excellence. Key Job Responsibilities Drive AWS's European backbone network infrastructure strategy and execution. Design and implement large scale network architecture supporting Amazon's global operations. Develop and execute strategic network expansion plans with regular milestone tracking. Optimize network topology for maximum efficiency and cost effectiveness. Monitor and analyze telecommunications market trends, emerging technologies, and industry developments. Identify and evaluate strategic suppliers for wavelengths, dark fiber, and submarine cable capacity. Lead technical vendor strategy across AWS's European network topology. Develop and maintain key commercial and technical partnerships within Europe. Manage supplier onboarding, performance metrics, and service quality standards. Coordinate with vendors and internal teams to deliver new backbone capacity. Collaborate with finance partners on budget planning and allocation. Manage risk assessment, escalations, and technical constraint resolution. Balance business requirements with technical limitations and cost considerations. Basic Qualifications Bachelor's degree in Information Technology, Computer Science, or a related field, or experience related IT. Experience in spectrum management, securing licensing for telecommunications services, and ensuring regulatory compliance at an international level. Experience managing procurement teams with direct experience in supplier/supply chain management. Experience in solving complex business challenges by delivering accurate and timely financial models, analysis, and recommendations that have a proven impact on business (e.g., financial savings, operational improvements, or customer benefits). Preferred Qualifications Experience in a network planning, construction/implementation, or commercial implementation role at an ISP, telecoms operator, or hyperscaler. Experience in strategic marketing management and market analysis, demonstrated ability to build and execute a strategy with clear goals and objectives aligned to business and service objectives. Experience working across functional teams and senior stakeholders. Experience with network protocols such as DNS, DHCP, TCP, or Linux and networking protocols, with strong analytical skills, attention to detail, and effective communication abilities. Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Hippo Digital Limited
Senior Test Engineer
Hippo Digital Limited
About The Role Hippo is a rapidly growing digital consultancy passionate about building and delivering transformative digital solutions. We are recruiting for a Senior Test Engineer to join our Hippo Herd. We work at the intersection of strategy, design, and technology to solve complex problems and create lasting value. Our collaborative and agile culture empowers our teams to make a genuine impact. As a Senior Test Engineer, you will play an important role in making Hippo the best Consultancy out there. You will work in multi-disciplinary teams that build, support, and maintain user-centred digital solutions. You will be responsible for analysing features, designing test parameters, creating customised quality checks, and writing final procedures. You will act as a Senior Consultant to deliver testing services to our clients. Senior Test Engineers at Hippo are experienced practitioners who lead technical deliverables, maintain client relationships, and are passionate about developing and upskilling others. Our solutions empower our customers to build and support secure, scalable, and well-engineered systems beyond traditional boundaries. We leverage deep data insights and continuous innovation to deliver awesome platforms that allow our customers to understand and get the most from their data and digital services. The Senior Test Engineer will be a key player and implementer in this mission. Please note we are looking for candidates who are looking for growth at the Senior level; therefore, the advertised salary band represents the entry point for this level, allowing for clear progression within the role, which we are very keen to support. Your Role in a Nutshell Meet with the product design team to determine testing parameters and write comprehensive test plans. Create test cases and conduct quality assurance and design performance tests using new procedures. Work collaboratively with colleagues to explore, design, and deliver solutions to client problems. Design and manage software development and deployment pipelines, resolving issues and potential bottlenecks. Present prototypes, solutions, and progress to internal and external stakeholders in a clear, concise manner. Build great relationships with your team and stakeholders, ensuring challenges are overcome. Lead in the recruitment, on-boarding, and line management of other engineers. Support other consultants in their professional development. Promote Hippo's Engineering Herd externally through blogs, workshops, or conferences. Skills and experience that you need Essential Experience Extensive experience in functional and non-functional automation testing. Extensive experience across the full testing lifecycle with a proven track record as a QA/Test Engineer. Skilled in testing web and mobile applications, including accessibility, cross-browser, and usability testing. Expert knowledge of testing tools such as Selenium, Specflow, Cucumber, Gherkin, JMeter, and K6. Experience with Containerisation tools like Docker or Kubernetes. Expertise in Microservices design and architecture. Solid experience with RESTful API Services. Mastery of Agile methods and techniques such as Scrum, Kanban, and TDD. Skilled in designing and building frameworks to test native and web applications across various scales. Experience with Infrastructure as Code (CloudFormation, Terraform). Expertise in Public Cloud (AWS or Azure) and DevOps. Proficient in CI/CD tools such as Jenkins, AWS CodeBuild, or Azure DevOps. Excellent verbal/written communication skills for technical and non-technical audiences. Desirable Experience Previous experience of working on GOV/NHS projects. Exposure to Python is highly desirable. What makes us great Alongside working with some of the most talented people out there and on-top of a competitive salary, which we're transparent about from the outset, you can also expect a range of benefits: Contributory Pension Scheme (Hippo 6% & Employee 2%) 25 Days Holiday plus UK Public Holidays Perkbox access for a wide range of discounts Critical illness cover Life assurance and death in service cover Volunteer days Cycle-to-work scheme for the avid cyclists Salary sacrifice electric vehicles scheme Season ticket loans Financial and general wellbeing sessions Flexible benefits scheme with options of: Private health cover Private dental cover Additional company pension contributions Additional holidays (up to an extra 2 days) Wellbeing contribution Charity contributions Tree planting Diversity, Inclusion and Belonging at Hippo At Hippo, we're dedicated to creating a diverse, equitable and inclusive workplace that works for everyone. We understand that having a diverse team unlocks our capacity for innovation, creativity and problem solving. Only by building a community of diverse perspectives, cultures and socio-economic backgrounds can we create an environment where all can contribute and thrive. We actively encourage applications from underrepresented groups including women, ethnic minorities, LGBTQ+, neurodivergent and people with disabilities. We are committed to providing an inclusive and accessible recruitment process that reflects our workplace culture. We are a registered Disability Confident Employer, Mindful Employer, Endometriosis Friendly Employer and a member of the Armed Forces Covenant. Hippo continually strives to remove barriers, provide accommodations and offer reasonable adjustments to ensure equity throughout our practices. Hi, we're Hippo At Hippo, we design with empathy and build for impact. We do this by combining data-informed evidence, human-centred design and software engineering. We're a digital services partner who is genuinely invested in helping our clients thrive as modern organisations. Our delivery methodology is truly agile, from concept to reality, supporting innovation and continuous improvement to achieve your desired outcomes. We firmly believe that technology should serve humanity, not the other way around. We take a human-centred approach to everything we do because we understand that complex problems require a service design approach. This means understanding how users behave and ensuring our solutions work for them in the real world. Our combination of data, design, and engineering delivers bespoke digital services that make a positive and meaningful impact on organisations and society. Have a look through our Website at Our Story & Purpose, Our Culture & Values, Case Studies and Our Solutions & Services to find out some more! Hippo Locations We are headquartered in Leeds and have offices across the UK in Glasgow, Manchester, Birmingham, London and Bristol. We're on the lookout for top talent nationwide but you need to be located within reasonable travelling distance from one of our offices. Given the dynamic nature of a consulting business, you may be required to work on-site at a Hippo office or at an in/out of town client location for a number of days per week (client dependent) and therefore candidates will need to be open/flexible to travel and working on one of those sites at least 2 days per week. We offer a generous relocation support package of up to £8,000 (please ask for terms and conditions) to help make your move a smooth one.
24/07/2026
Full time
About The Role Hippo is a rapidly growing digital consultancy passionate about building and delivering transformative digital solutions. We are recruiting for a Senior Test Engineer to join our Hippo Herd. We work at the intersection of strategy, design, and technology to solve complex problems and create lasting value. Our collaborative and agile culture empowers our teams to make a genuine impact. As a Senior Test Engineer, you will play an important role in making Hippo the best Consultancy out there. You will work in multi-disciplinary teams that build, support, and maintain user-centred digital solutions. You will be responsible for analysing features, designing test parameters, creating customised quality checks, and writing final procedures. You will act as a Senior Consultant to deliver testing services to our clients. Senior Test Engineers at Hippo are experienced practitioners who lead technical deliverables, maintain client relationships, and are passionate about developing and upskilling others. Our solutions empower our customers to build and support secure, scalable, and well-engineered systems beyond traditional boundaries. We leverage deep data insights and continuous innovation to deliver awesome platforms that allow our customers to understand and get the most from their data and digital services. The Senior Test Engineer will be a key player and implementer in this mission. Please note we are looking for candidates who are looking for growth at the Senior level; therefore, the advertised salary band represents the entry point for this level, allowing for clear progression within the role, which we are very keen to support. Your Role in a Nutshell Meet with the product design team to determine testing parameters and write comprehensive test plans. Create test cases and conduct quality assurance and design performance tests using new procedures. Work collaboratively with colleagues to explore, design, and deliver solutions to client problems. Design and manage software development and deployment pipelines, resolving issues and potential bottlenecks. Present prototypes, solutions, and progress to internal and external stakeholders in a clear, concise manner. Build great relationships with your team and stakeholders, ensuring challenges are overcome. Lead in the recruitment, on-boarding, and line management of other engineers. Support other consultants in their professional development. Promote Hippo's Engineering Herd externally through blogs, workshops, or conferences. Skills and experience that you need Essential Experience Extensive experience in functional and non-functional automation testing. Extensive experience across the full testing lifecycle with a proven track record as a QA/Test Engineer. Skilled in testing web and mobile applications, including accessibility, cross-browser, and usability testing. Expert knowledge of testing tools such as Selenium, Specflow, Cucumber, Gherkin, JMeter, and K6. Experience with Containerisation tools like Docker or Kubernetes. Expertise in Microservices design and architecture. Solid experience with RESTful API Services. Mastery of Agile methods and techniques such as Scrum, Kanban, and TDD. Skilled in designing and building frameworks to test native and web applications across various scales. Experience with Infrastructure as Code (CloudFormation, Terraform). Expertise in Public Cloud (AWS or Azure) and DevOps. Proficient in CI/CD tools such as Jenkins, AWS CodeBuild, or Azure DevOps. Excellent verbal/written communication skills for technical and non-technical audiences. Desirable Experience Previous experience of working on GOV/NHS projects. Exposure to Python is highly desirable. What makes us great Alongside working with some of the most talented people out there and on-top of a competitive salary, which we're transparent about from the outset, you can also expect a range of benefits: Contributory Pension Scheme (Hippo 6% & Employee 2%) 25 Days Holiday plus UK Public Holidays Perkbox access for a wide range of discounts Critical illness cover Life assurance and death in service cover Volunteer days Cycle-to-work scheme for the avid cyclists Salary sacrifice electric vehicles scheme Season ticket loans Financial and general wellbeing sessions Flexible benefits scheme with options of: Private health cover Private dental cover Additional company pension contributions Additional holidays (up to an extra 2 days) Wellbeing contribution Charity contributions Tree planting Diversity, Inclusion and Belonging at Hippo At Hippo, we're dedicated to creating a diverse, equitable and inclusive workplace that works for everyone. We understand that having a diverse team unlocks our capacity for innovation, creativity and problem solving. Only by building a community of diverse perspectives, cultures and socio-economic backgrounds can we create an environment where all can contribute and thrive. We actively encourage applications from underrepresented groups including women, ethnic minorities, LGBTQ+, neurodivergent and people with disabilities. We are committed to providing an inclusive and accessible recruitment process that reflects our workplace culture. We are a registered Disability Confident Employer, Mindful Employer, Endometriosis Friendly Employer and a member of the Armed Forces Covenant. Hippo continually strives to remove barriers, provide accommodations and offer reasonable adjustments to ensure equity throughout our practices. Hi, we're Hippo At Hippo, we design with empathy and build for impact. We do this by combining data-informed evidence, human-centred design and software engineering. We're a digital services partner who is genuinely invested in helping our clients thrive as modern organisations. Our delivery methodology is truly agile, from concept to reality, supporting innovation and continuous improvement to achieve your desired outcomes. We firmly believe that technology should serve humanity, not the other way around. We take a human-centred approach to everything we do because we understand that complex problems require a service design approach. This means understanding how users behave and ensuring our solutions work for them in the real world. Our combination of data, design, and engineering delivers bespoke digital services that make a positive and meaningful impact on organisations and society. Have a look through our Website at Our Story & Purpose, Our Culture & Values, Case Studies and Our Solutions & Services to find out some more! Hippo Locations We are headquartered in Leeds and have offices across the UK in Glasgow, Manchester, Birmingham, London and Bristol. We're on the lookout for top talent nationwide but you need to be located within reasonable travelling distance from one of our offices. Given the dynamic nature of a consulting business, you may be required to work on-site at a Hippo office or at an in/out of town client location for a number of days per week (client dependent) and therefore candidates will need to be open/flexible to travel and working on one of those sites at least 2 days per week. We offer a generous relocation support package of up to £8,000 (please ask for terms and conditions) to help make your move a smooth one.
Platform Chapter Lead - Engineering
easyJet Airline Company PLC
Platform Chapter Lead - Engineering We are easyJet - a FTSE-250 listed, £multi-billion low-cost airline that serves tens of millions of customers every single year. If you're reading this, you have probably already been an easyJet customer, and you'll know that there is no more iconic (or Orange!) travel brand in Europe. We fly more than 1,207 routes, connecting 38 countries across Europe, and employ more than 18,000 colleagues. We're on a mission to make low-cost travel easy - and whatever your role here, you'll connect millions of people to what they love using Europe's best airline network, great value fares, and friendly service. What makes us easyJet? Our Promise Behaviours - we are Safe, Bold, Welcoming and Challenging. Four Behaviours. One Spirit. One easyJet. About the role We're modernising how we build and ship software across our digital and operational estate. As Head of Platform Engineering you'll own the strategy, delivery and operation of our internal developer platform, and lead the group of teams behind it. Platform engineering at an airline is a genuinely interesting problem: revenue critical services that have to stay up through seasonal peaks and disruption driven traffic spikes, a large and evolving estate, and real pace. You'll set direction, lead through your managers, and own a budget - and while you won't be writing code, you'll be close enough to challenge technical choices and hold your own in a design debate. Why now. We're investing in platform engineering as a core capability, and this is a pivotal hire. There's genuine scope to shape the teams, define how the platform is built and run, and raise the engineering bar - not to maintain a finished thing, but to build it into something great. Reports to: Director of Engineering Direct reports: 4-5 team leads / engineering managers (Platform & DevOps) What you'll do Lead the teams. Lead 4-5 Platform, DevOps and Developer Enablement teams through their managers; develop your leaders, plan capacity, and hire and grow strong talent. Own the strategy. Define and drive the roadmap for our Internal Developer Platform and tooling, set technical direction with Enterprise Architecture, and make deliberate build/buy calls. Own delivery. Hold end to end accountability for delivery - quality, predictability and pace - surfacing risks early and driving continuous improvement. Own the budget. Own the operating budget (opex) and headcount, run a cost effective tooling strategy, and embed FinOps to manage cloud spend and demonstrate value. Drive developer experience. Make the platform something teams want to use - paved paths, less friction - using metrics (e.g. DORA) and real feedback to keep improving it. Keep it reliable and compliant. Oversee performance, resilience and observability for revenue critical services through peak traffic, and maintain security and compliance (e.g. PCI DSS, GDPR). Bring the business with you. Align platform strategy with business goals and translate technical trade offs into clear decisions, up to executive level. What you'll bring You live and breathe modern platform engineering - Continuous Delivery, Team Topologies, developer experience and the craft of building strong engineering culture. These aren't buzzwords to you; they're how you think. A track record of leading multiple engineering teams at scale (as a leader of managers/leads), setting platform strategy and turning it into delivered outcomes. Ownership of budget and headcount, with the commercial judgement to balance cost, speed and risk. Strong cloud native and modern DevOps background - cloud (ideally AWS) and Kubernetes, CI/CD, infrastructure as code and observability - with enough engineering depth (e.g. Java, .NET, Python) to be credible with strong engineers. Excellent communication and the ability to influence at executive level, plus a real record of growing engineering leaders. Building or operating an Internal Developer Platform or developer tooling - Backstage especially welcome. Familiarity with tools like Terraform, GitHub Actions and Octopus (or equivalents), and FinOps / cloud cost management. Regulated, high availability or high transaction volume domains (travel, retail, payments). What success looks like in the first 12 months A clear platform roadmap and operating model the teams are bought into. Measurably improved delivery predictability and developer experience. A healthy fin ops culture embedded and a Product led leadership team supporting you. Please note that this role does not meet the criteria for visa sponsorship, and we are therefore unable to consider applicants who require sponsorship to work in the UK. Benefits Competitive base salary 25 days holiday, pension scheme, life assurance, and a flexible benefits package Discounted staff travel scheme for friends and family Annual credit for discount on easyJet holidays Work Away scheme, allowing you to work abroad for 30 days a year Electric vehicle lease salary sacrifice scheme Location & Hours of Work We operate a hybrid working policy of 40%-60% of the month spent with colleagues.
23/07/2026
Full time
Platform Chapter Lead - Engineering We are easyJet - a FTSE-250 listed, £multi-billion low-cost airline that serves tens of millions of customers every single year. If you're reading this, you have probably already been an easyJet customer, and you'll know that there is no more iconic (or Orange!) travel brand in Europe. We fly more than 1,207 routes, connecting 38 countries across Europe, and employ more than 18,000 colleagues. We're on a mission to make low-cost travel easy - and whatever your role here, you'll connect millions of people to what they love using Europe's best airline network, great value fares, and friendly service. What makes us easyJet? Our Promise Behaviours - we are Safe, Bold, Welcoming and Challenging. Four Behaviours. One Spirit. One easyJet. About the role We're modernising how we build and ship software across our digital and operational estate. As Head of Platform Engineering you'll own the strategy, delivery and operation of our internal developer platform, and lead the group of teams behind it. Platform engineering at an airline is a genuinely interesting problem: revenue critical services that have to stay up through seasonal peaks and disruption driven traffic spikes, a large and evolving estate, and real pace. You'll set direction, lead through your managers, and own a budget - and while you won't be writing code, you'll be close enough to challenge technical choices and hold your own in a design debate. Why now. We're investing in platform engineering as a core capability, and this is a pivotal hire. There's genuine scope to shape the teams, define how the platform is built and run, and raise the engineering bar - not to maintain a finished thing, but to build it into something great. Reports to: Director of Engineering Direct reports: 4-5 team leads / engineering managers (Platform & DevOps) What you'll do Lead the teams. Lead 4-5 Platform, DevOps and Developer Enablement teams through their managers; develop your leaders, plan capacity, and hire and grow strong talent. Own the strategy. Define and drive the roadmap for our Internal Developer Platform and tooling, set technical direction with Enterprise Architecture, and make deliberate build/buy calls. Own delivery. Hold end to end accountability for delivery - quality, predictability and pace - surfacing risks early and driving continuous improvement. Own the budget. Own the operating budget (opex) and headcount, run a cost effective tooling strategy, and embed FinOps to manage cloud spend and demonstrate value. Drive developer experience. Make the platform something teams want to use - paved paths, less friction - using metrics (e.g. DORA) and real feedback to keep improving it. Keep it reliable and compliant. Oversee performance, resilience and observability for revenue critical services through peak traffic, and maintain security and compliance (e.g. PCI DSS, GDPR). Bring the business with you. Align platform strategy with business goals and translate technical trade offs into clear decisions, up to executive level. What you'll bring You live and breathe modern platform engineering - Continuous Delivery, Team Topologies, developer experience and the craft of building strong engineering culture. These aren't buzzwords to you; they're how you think. A track record of leading multiple engineering teams at scale (as a leader of managers/leads), setting platform strategy and turning it into delivered outcomes. Ownership of budget and headcount, with the commercial judgement to balance cost, speed and risk. Strong cloud native and modern DevOps background - cloud (ideally AWS) and Kubernetes, CI/CD, infrastructure as code and observability - with enough engineering depth (e.g. Java, .NET, Python) to be credible with strong engineers. Excellent communication and the ability to influence at executive level, plus a real record of growing engineering leaders. Building or operating an Internal Developer Platform or developer tooling - Backstage especially welcome. Familiarity with tools like Terraform, GitHub Actions and Octopus (or equivalents), and FinOps / cloud cost management. Regulated, high availability or high transaction volume domains (travel, retail, payments). What success looks like in the first 12 months A clear platform roadmap and operating model the teams are bought into. Measurably improved delivery predictability and developer experience. A healthy fin ops culture embedded and a Product led leadership team supporting you. Please note that this role does not meet the criteria for visa sponsorship, and we are therefore unable to consider applicants who require sponsorship to work in the UK. Benefits Competitive base salary 25 days holiday, pension scheme, life assurance, and a flexible benefits package Discounted staff travel scheme for friends and family Annual credit for discount on easyJet holidays Work Away scheme, allowing you to work abroad for 30 days a year Electric vehicle lease salary sacrifice scheme Location & Hours of Work We operate a hybrid working policy of 40%-60% of the month spent with colleagues.
Senior Cloud SRE - AI/ML Platform & GPU Compute
Icehouseventures
At Wayve we're committed to creating a diverse, fair and respectful culture that is inclusive of everyone based on their unique skills and perspectives, and regardless of sex, race, religion or belief, ethnic or national origin, disability, age, citizenship, marital, domestic or civil partnership status, sexual orientation, gender identity, veteran status, pregnancy or related condition (including breastfeeding) or any other basis as protected by applicable law. About us Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing the usability and safety of automated driving systems. Our vision is to create autonomy that propels the world forward. Our intelligent, mapless, and hardware-agnostic AI products are designed for automakers, accelerating the transition from assisted to automated driving. In our fast-paced environment big problems ignite us-we embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future. At Wayve, your contributions matter. We value diversity, embrace new perspectives, and foster an inclusive work environment; we back each other to deliver impact. Make Wayve the experience that defines your career! The role As a Cloud Site Reliability Engineer at Wayve, you will build and scale the reliability foundations of our AI cloud platform. This includes our Model Development Platform (powering end-to-end model development from raw data to on-road experimentation) and our GPU Compute platform (large-scale, multi-tenant GPU fleets and scheduling systems driving model training and inference at scale). This is a founding Cloud SRE role. You won't inherit a mature SRE function, you'll help create it. You will define the frameworks, automation, and operational standards that ensure our model development infrastructure, distributed systems, and large compute clusters operate predictably, efficiently, and at scale. This role sits at the intersection of AI research, large-scale cloud infrastructure, and production operations. Your work will directly enable faster model training, reliable experimentation, and scalable AI deployment by ensuring our cloud infrastructure is resilient and performant. Key responsibilities Reliability & Platform Ownership Own the reliability, availability, and performance of the Model Dev Platform and GPU Compute environments. Define and operationalise SLOs, SLIs, and error budgets across platform services. Improve capacity planning, scaling strategies, and resource efficiency across large GPU-backed clusters. Partner with ML, platform, and software teams to establish clear production readiness standards. Incident Response & On-Call Participate in a 24/7 on-call rotation as first-line response for cloud and cluster-related incidents. Lead incident triage, escalation, communications, and root cause analysis. Translate post-incident learning into durable architectural or automation improvements. Continuously reduce alert noise and recurring operational burden. Observability & Operational Excellence Design and operate monitoring, logging, tracing, and alerting systems that enable rapid detection and recovery. Build dashboards that reflect real user-centric platform health (not just infrastructure metrics). Improve deployment safety through better change management, validation, and rollback mechanisms. Automation & Tooling Build automation for cluster operations, training workflows, remediation, and scaling tasks. Implement self healing patterns and resilient recovery workflows. Harden CI/CD and release processes to improve deployment safety and velocity. Support infrastructure as code and policy driven guardrails to ensure secure, reliable cloud environments. About you In order to set you up for success as a Cloud Site Reliability Engineer at Wayve, we're looking for the following skills and experience. Essential skills Proven experience in an SRE, Production Engineer, or Cloud Reliability role supporting large-scale cloud systems. Strong Kubernetes experience, including operating production clusters. Hands on experience running production workloads in AWS, GCP, or Azure. Experience operating complex distributed systems in production, ideally including compute-heavy or high-performance workloads. Experience working with large compute clusters; exposure to AI/ML training or inference workloads strongly preferred. Strong Linux fundamentals and proficiency in at least one scripting or systems language (e.g., Python, Go, C++) with a bias toward automation. Deep troubleshooting skills across networking, storage, distributed systems, and performance at scale. Experience designing and operating observability stacks (e.g., Datadog, Prometheus, Grafana, OpenTelemetry). Clear communication skills, including leading incidents, writing post mortems, and influencing teams to prioritise reliability improvements. Desirable skills Experience operating GPU backed environments or large scale ML infrastructure. Experience running model training or inference pipelines in production (MLOps). Familiarity with infrastructure as code (e.g., Terraform) and secure cloud production environments. Experience defining and running SLOs/SLIs and building reliability programs across multiple teams. Experience as an early or founding SRE hire establishing processes from scratch. Interest in helping shape and grow a Cloud SRE function, with potential to take on leadership responsibilities over time. This is a full time role based in our office in London (2 days a week in the office). At Wayve we want the best of all worlds so we operate a hybrid working policy that combines time together in our offices and workshops to fuel innovation, culture, relationships and learning, and time spent working from home. Wayve is committed to creating an inclusive interview experience. If you require any accommodations or adjustments to participate fully in our interview process, please let us know.
22/07/2026
Full time
At Wayve we're committed to creating a diverse, fair and respectful culture that is inclusive of everyone based on their unique skills and perspectives, and regardless of sex, race, religion or belief, ethnic or national origin, disability, age, citizenship, marital, domestic or civil partnership status, sexual orientation, gender identity, veteran status, pregnancy or related condition (including breastfeeding) or any other basis as protected by applicable law. About us Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing the usability and safety of automated driving systems. Our vision is to create autonomy that propels the world forward. Our intelligent, mapless, and hardware-agnostic AI products are designed for automakers, accelerating the transition from assisted to automated driving. In our fast-paced environment big problems ignite us-we embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future. At Wayve, your contributions matter. We value diversity, embrace new perspectives, and foster an inclusive work environment; we back each other to deliver impact. Make Wayve the experience that defines your career! The role As a Cloud Site Reliability Engineer at Wayve, you will build and scale the reliability foundations of our AI cloud platform. This includes our Model Development Platform (powering end-to-end model development from raw data to on-road experimentation) and our GPU Compute platform (large-scale, multi-tenant GPU fleets and scheduling systems driving model training and inference at scale). This is a founding Cloud SRE role. You won't inherit a mature SRE function, you'll help create it. You will define the frameworks, automation, and operational standards that ensure our model development infrastructure, distributed systems, and large compute clusters operate predictably, efficiently, and at scale. This role sits at the intersection of AI research, large-scale cloud infrastructure, and production operations. Your work will directly enable faster model training, reliable experimentation, and scalable AI deployment by ensuring our cloud infrastructure is resilient and performant. Key responsibilities Reliability & Platform Ownership Own the reliability, availability, and performance of the Model Dev Platform and GPU Compute environments. Define and operationalise SLOs, SLIs, and error budgets across platform services. Improve capacity planning, scaling strategies, and resource efficiency across large GPU-backed clusters. Partner with ML, platform, and software teams to establish clear production readiness standards. Incident Response & On-Call Participate in a 24/7 on-call rotation as first-line response for cloud and cluster-related incidents. Lead incident triage, escalation, communications, and root cause analysis. Translate post-incident learning into durable architectural or automation improvements. Continuously reduce alert noise and recurring operational burden. Observability & Operational Excellence Design and operate monitoring, logging, tracing, and alerting systems that enable rapid detection and recovery. Build dashboards that reflect real user-centric platform health (not just infrastructure metrics). Improve deployment safety through better change management, validation, and rollback mechanisms. Automation & Tooling Build automation for cluster operations, training workflows, remediation, and scaling tasks. Implement self healing patterns and resilient recovery workflows. Harden CI/CD and release processes to improve deployment safety and velocity. Support infrastructure as code and policy driven guardrails to ensure secure, reliable cloud environments. About you In order to set you up for success as a Cloud Site Reliability Engineer at Wayve, we're looking for the following skills and experience. Essential skills Proven experience in an SRE, Production Engineer, or Cloud Reliability role supporting large-scale cloud systems. Strong Kubernetes experience, including operating production clusters. Hands on experience running production workloads in AWS, GCP, or Azure. Experience operating complex distributed systems in production, ideally including compute-heavy or high-performance workloads. Experience working with large compute clusters; exposure to AI/ML training or inference workloads strongly preferred. Strong Linux fundamentals and proficiency in at least one scripting or systems language (e.g., Python, Go, C++) with a bias toward automation. Deep troubleshooting skills across networking, storage, distributed systems, and performance at scale. Experience designing and operating observability stacks (e.g., Datadog, Prometheus, Grafana, OpenTelemetry). Clear communication skills, including leading incidents, writing post mortems, and influencing teams to prioritise reliability improvements. Desirable skills Experience operating GPU backed environments or large scale ML infrastructure. Experience running model training or inference pipelines in production (MLOps). Familiarity with infrastructure as code (e.g., Terraform) and secure cloud production environments. Experience defining and running SLOs/SLIs and building reliability programs across multiple teams. Experience as an early or founding SRE hire establishing processes from scratch. Interest in helping shape and grow a Cloud SRE function, with potential to take on leadership responsibilities over time. This is a full time role based in our office in London (2 days a week in the office). At Wayve we want the best of all worlds so we operate a hybrid working policy that combines time together in our offices and workshops to fuel innovation, culture, relationships and learning, and time spent working from home. Wayve is committed to creating an inclusive interview experience. If you require any accommodations or adjustments to participate fully in our interview process, please let us know.
Lead Site Reliability Engineer
United States Digital Space LLC
JOB DESCRIPTION Our trading technology stack is undergoing a multi year convergence and modernization journey. You will play a pivotal role in shaping our next generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast paced front office environments, enjoys direct interaction with traders, and wants to influence the reliability culture of a major global trading organization. As a Lead Site Reliability Engineer at JPMorgan Chase within J.P. Morgan Asset Management's Trading Technology group, you will be embedded directly within the software engineering team that builds and supports our front office trading platforms. Job Responsibilities Engage daily with traders across asset classes (Equities, Fixed Income, FX) to understand workflows, pain points, and reliability priorities. Act as a trusted engineering partner to the desk, ensuring systems are stable, performant, and aligned with business needs. Support live trading environments, including incident response, root cause analysis, and post mortem leadership. Work as a core member of the software engineering team, participating in daily standups and design discussions. Contribute directly to the codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation. Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self healing workflows, and resilience engineering. Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements. Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls. Drive improvements in latency, throughput, and stability across high volume trading applications. Build and maintain tooling for monitoring, alerting, and distributed tracing across global environments. Operate within a globally distributed engineering and trading organization, collaborating with teams in EMEA, US, and APAC. Partner with infrastructure, networking, cloud engineering, and cybersecurity teams to ensure end to end reliability. Required Qualifications, Capabilities, and Skills Strong hands on experience in front office trading environments or similarly high pressure, low latency domains. Proficiency with SRE tooling and techniques, including FIX messaging, Kafka, Grafana, Splunk, ITRS Geneos, Dynatrace, InfluxDB, MQ (IBM MQ or similar), Oracle DB Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity. Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations. Deep knowledge of reliability engineering principles: SLIs/SLOs, real-time telemetry, disaster recovery planning, capacity planning, and performance tuning. Experience designing and implementing observability frameworks for mission critical systems. Proven ability to lead incident response and drive long term remediation. Solid programming skills in Python, Java, or Kotlin, with the ability to contribute production grade code. Experience with microservices, distributed systems, and event driven architectures. Strong understanding of CI/CD pipelines, automated testing, and deployment strategies. Comfortable interacting directly with traders and senior stakeholders. Excellent communication skills, especially when translating technical issues into business impact. Ability to operate calmly and decisively in high pressure situations. Strong leadership presence with a collaborative mindset. ABOUT US J.P. Morgan is a global leader in financial services, providing strategic advice and products to the world's most prominent corporations, governments, wealthy individuals and institutional investors. Our first-class business in a first-class way approach to serving clients drives everything we do. We strive to build trusted, long-term partnerships to help our clients achieve their business objectives. We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants' and employees' religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation. ABOUT THE TEAM J.P. Morgan Asset & Wealth Management delivers industry-leading investment management and private banking solutions. Asset Management provides individuals, advisors and institutions with strategies and expertise that span the full spectrum of asset classes through our global network of investment professionals. Wealth Management helps individuals, families and foundations take a more intentional approach to their wealth or finances to better define, focus and realize their goals.
22/07/2026
Full time
JOB DESCRIPTION Our trading technology stack is undergoing a multi year convergence and modernization journey. You will play a pivotal role in shaping our next generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast paced front office environments, enjoys direct interaction with traders, and wants to influence the reliability culture of a major global trading organization. As a Lead Site Reliability Engineer at JPMorgan Chase within J.P. Morgan Asset Management's Trading Technology group, you will be embedded directly within the software engineering team that builds and supports our front office trading platforms. Job Responsibilities Engage daily with traders across asset classes (Equities, Fixed Income, FX) to understand workflows, pain points, and reliability priorities. Act as a trusted engineering partner to the desk, ensuring systems are stable, performant, and aligned with business needs. Support live trading environments, including incident response, root cause analysis, and post mortem leadership. Work as a core member of the software engineering team, participating in daily standups and design discussions. Contribute directly to the codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation. Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self healing workflows, and resilience engineering. Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements. Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls. Drive improvements in latency, throughput, and stability across high volume trading applications. Build and maintain tooling for monitoring, alerting, and distributed tracing across global environments. Operate within a globally distributed engineering and trading organization, collaborating with teams in EMEA, US, and APAC. Partner with infrastructure, networking, cloud engineering, and cybersecurity teams to ensure end to end reliability. Required Qualifications, Capabilities, and Skills Strong hands on experience in front office trading environments or similarly high pressure, low latency domains. Proficiency with SRE tooling and techniques, including FIX messaging, Kafka, Grafana, Splunk, ITRS Geneos, Dynatrace, InfluxDB, MQ (IBM MQ or similar), Oracle DB Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity. Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations. Deep knowledge of reliability engineering principles: SLIs/SLOs, real-time telemetry, disaster recovery planning, capacity planning, and performance tuning. Experience designing and implementing observability frameworks for mission critical systems. Proven ability to lead incident response and drive long term remediation. Solid programming skills in Python, Java, or Kotlin, with the ability to contribute production grade code. Experience with microservices, distributed systems, and event driven architectures. Strong understanding of CI/CD pipelines, automated testing, and deployment strategies. Comfortable interacting directly with traders and senior stakeholders. Excellent communication skills, especially when translating technical issues into business impact. Ability to operate calmly and decisively in high pressure situations. Strong leadership presence with a collaborative mindset. ABOUT US J.P. Morgan is a global leader in financial services, providing strategic advice and products to the world's most prominent corporations, governments, wealthy individuals and institutional investors. Our first-class business in a first-class way approach to serving clients drives everything we do. We strive to build trusted, long-term partnerships to help our clients achieve their business objectives. We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants' and employees' religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation. ABOUT THE TEAM J.P. Morgan Asset & Wealth Management delivers industry-leading investment management and private banking solutions. Asset Management provides individuals, advisors and institutions with strategies and expertise that span the full spectrum of asset classes through our global network of investment professionals. Wealth Management helps individuals, families and foundations take a more intentional approach to their wealth or finances to better define, focus and realize their goals.
AND Digital
Observability Architect - 12 Month FTC
AND Digital
Observability Architect 12 Month FTC About you: You care deeply about producing high-quality work that delivers real value You're comfortable navigating ambiguity and solving complex problems collaboratively You bring strong expertise in your craft, alongside a willingness to keep learning You communicate clearly and build trust quickly with clients and teammates You're pragmatic, adaptable and outcome-focused You enjoy sharing knowledge and helping others grow You value low-ego collaboration and enjoy working as part of multidisciplinary teams Role Objective Lead the assessment, design, and optimisation of the observability strategy for the co-location migration programme. Ensure logging, metrics, tracing, alerting, and operational dashboards provide comprehensive visibility across the new infrastructure and application estate, enabling the successful migration of the tightly-coupled monolithic platform with minimal operational risk. Identify gaps in the existing observability capability and recommend enhancements to tooling, processes, and architecture where required. Key Responsibilities Observability Assessment & Strategy Review the current observability architecture across infrastructure, networks, middleware, databases, and applications. Assess existing logging, metrics, distributed tracing, and monitoring capabilities to determine readiness for the co-location migration. Develop an observability strategy that supports both migration activities and long-term operational support. Recommend enhancements or platform uplifts where current tooling does not provide sufficient visibility or resilience. Baseline Performance Analysis Analyse telemetry, monitoring data, dashboards, and operational trends from completed migration waves. Establish performance baselines for compute, storage, networking, application response times, and transaction throughput. Identify recurring operational issues and use historical insights to improve migration readiness. Define measurable service health indicators to compare pre- and post-migration performance. Monolithic Application Monitoring Design comprehensive monitoring for the tightly-coupled monolithic application estate, with particular emphasis on latency-sensitive interdependencies. Create real-time dashboards that provide operational visibility across infrastructure, middleware, databases, messaging, and application components. Ensure end-to-end transaction tracing is available to rapidly identify bottlenecks and service degradation. Validate monitoring coverage prior to each migration wave. Logging & Trace Management Review and standardise centralised logging across all migrated environments. Ensure consistent log formats, metadata, correlation IDs, and traceability across systems. Validate log ingestion, retention policies, indexing, and search performance. Ensure operational teams can rapidly investigate incidents using correlated logs and distributed traces. Alerting & Operational Readiness Review and optimise alert thresholds to minimise both missed events and unnecessary alert noise. Implement intelligent alerting aligned to business services and critical customer journeys. Define migration-specific alerting for infrastructure failures, application degradation, latency increases, replication issues, and capacity constraints. Support operational readiness activities including rehearsals and production cutover monitoring. Compliance & Audit Ensure observability solutions meet financial services regulatory requirements for auditability, log retention, security, and data governance. Validate access controls and security monitoring for observability platforms. Support evidence gathering for internal governance, audit, and regulatory reviews. Platform Improvement Evaluate the suitability of existing observability platforms and recommend improvements where required. Assess opportunities to improve automation, anomaly detection, service health monitoring, and predictive alerting. Define standards and best practices for observability across future migration phases. Stakeholder Collaboration Work closely with Infrastructure Architects, Application Architects, Platform Engineering, Security, Operations, and Migration teams. Provide technical guidance during migration planning, testing, dress rehearsals, and production cutovers. Produce architecture documentation, monitoring standards, operational runbooks, and knowledge transfer materials. Required Skills & Experience Extensive experience designing enterprise observability solutions within large scale infrastructure or data centre migration programmes. Strong knowledge of metrics, logging, distributed tracing, and application performance monitoring (APM). Experience monitoring latency sensitive, business critical enterprise applications. Strong understanding of infrastructure, virtualisation, networking, storage, databases, and middleware monitoring. Experience implementing centralised logging and observability best practices. Knowledge of financial services operational resilience, audit, and regulatory requirements. Ability to analyse complex operational telemetry and identify performance bottlenecks. Excellent stakeholder management and communication skills. By joining AND, we'll provide: 25 days bookable holiday + flexible Bank Holidays Pension: 6% of salary paid by AND Digital with a further 2% paid by you (can be increased by choice). Aviva healthcare cover (including pre existing condition cover) for you. Flexibenefit: £1000 assigned to you via our benefits portal to select or upgrade the benefits that suit you the most. Any unused allowance from the £1000 can be taken as cash. Life Assurance. Income Protection. Eye test + first pair of glasses. Parental Benefits: Generous Enhanced Maternity and Enhanced Partner (Paternity) Leave. Equal Opportunities Statement Diversity and inclusion are hugely important to us, and we're committed to providing equal opportunities for all. We're actively recruiting for a diverse and inclusive workforce so want to ensure we do everything we can to support your application. We want you to feel safe and empowered to let us know if you need any adjustments to be made to your application or interview process, so please speak to our recruitment team.
21/07/2026
Full time
Observability Architect 12 Month FTC About you: You care deeply about producing high-quality work that delivers real value You're comfortable navigating ambiguity and solving complex problems collaboratively You bring strong expertise in your craft, alongside a willingness to keep learning You communicate clearly and build trust quickly with clients and teammates You're pragmatic, adaptable and outcome-focused You enjoy sharing knowledge and helping others grow You value low-ego collaboration and enjoy working as part of multidisciplinary teams Role Objective Lead the assessment, design, and optimisation of the observability strategy for the co-location migration programme. Ensure logging, metrics, tracing, alerting, and operational dashboards provide comprehensive visibility across the new infrastructure and application estate, enabling the successful migration of the tightly-coupled monolithic platform with minimal operational risk. Identify gaps in the existing observability capability and recommend enhancements to tooling, processes, and architecture where required. Key Responsibilities Observability Assessment & Strategy Review the current observability architecture across infrastructure, networks, middleware, databases, and applications. Assess existing logging, metrics, distributed tracing, and monitoring capabilities to determine readiness for the co-location migration. Develop an observability strategy that supports both migration activities and long-term operational support. Recommend enhancements or platform uplifts where current tooling does not provide sufficient visibility or resilience. Baseline Performance Analysis Analyse telemetry, monitoring data, dashboards, and operational trends from completed migration waves. Establish performance baselines for compute, storage, networking, application response times, and transaction throughput. Identify recurring operational issues and use historical insights to improve migration readiness. Define measurable service health indicators to compare pre- and post-migration performance. Monolithic Application Monitoring Design comprehensive monitoring for the tightly-coupled monolithic application estate, with particular emphasis on latency-sensitive interdependencies. Create real-time dashboards that provide operational visibility across infrastructure, middleware, databases, messaging, and application components. Ensure end-to-end transaction tracing is available to rapidly identify bottlenecks and service degradation. Validate monitoring coverage prior to each migration wave. Logging & Trace Management Review and standardise centralised logging across all migrated environments. Ensure consistent log formats, metadata, correlation IDs, and traceability across systems. Validate log ingestion, retention policies, indexing, and search performance. Ensure operational teams can rapidly investigate incidents using correlated logs and distributed traces. Alerting & Operational Readiness Review and optimise alert thresholds to minimise both missed events and unnecessary alert noise. Implement intelligent alerting aligned to business services and critical customer journeys. Define migration-specific alerting for infrastructure failures, application degradation, latency increases, replication issues, and capacity constraints. Support operational readiness activities including rehearsals and production cutover monitoring. Compliance & Audit Ensure observability solutions meet financial services regulatory requirements for auditability, log retention, security, and data governance. Validate access controls and security monitoring for observability platforms. Support evidence gathering for internal governance, audit, and regulatory reviews. Platform Improvement Evaluate the suitability of existing observability platforms and recommend improvements where required. Assess opportunities to improve automation, anomaly detection, service health monitoring, and predictive alerting. Define standards and best practices for observability across future migration phases. Stakeholder Collaboration Work closely with Infrastructure Architects, Application Architects, Platform Engineering, Security, Operations, and Migration teams. Provide technical guidance during migration planning, testing, dress rehearsals, and production cutovers. Produce architecture documentation, monitoring standards, operational runbooks, and knowledge transfer materials. Required Skills & Experience Extensive experience designing enterprise observability solutions within large scale infrastructure or data centre migration programmes. Strong knowledge of metrics, logging, distributed tracing, and application performance monitoring (APM). Experience monitoring latency sensitive, business critical enterprise applications. Strong understanding of infrastructure, virtualisation, networking, storage, databases, and middleware monitoring. Experience implementing centralised logging and observability best practices. Knowledge of financial services operational resilience, audit, and regulatory requirements. Ability to analyse complex operational telemetry and identify performance bottlenecks. Excellent stakeholder management and communication skills. By joining AND, we'll provide: 25 days bookable holiday + flexible Bank Holidays Pension: 6% of salary paid by AND Digital with a further 2% paid by you (can be increased by choice). Aviva healthcare cover (including pre existing condition cover) for you. Flexibenefit: £1000 assigned to you via our benefits portal to select or upgrade the benefits that suit you the most. Any unused allowance from the £1000 can be taken as cash. Life Assurance. Income Protection. Eye test + first pair of glasses. Parental Benefits: Generous Enhanced Maternity and Enhanced Partner (Paternity) Leave. Equal Opportunities Statement Diversity and inclusion are hugely important to us, and we're committed to providing equal opportunities for all. We're actively recruiting for a diverse and inclusive workforce so want to ensure we do everything we can to support your application. We want you to feel safe and empowered to let us know if you need any adjustments to be made to your application or interview process, so please speak to our recruitment team.
Lead Site Reliability Engineer
JPMorgan Chase & Co.
Our trading technology stack is undergoing a multi year convergence and modernization journey. You will play a pivotal role in shaping our next generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast paced front office environments, enjoys direct interaction with traders, and wants to influence the reliability culture of a major global trading organization. As a Lead Site Reliability Engineer at JPMorgan Chase within J.P. Morgan Asset Management's Trading Technology group, you will be embedded directly within the software engineering team that builds and supports our front office trading platforms. Job Responsibilities: Engage daily with traders across asset classes (Equities, Fixed Income, FX) to understand workflows, pain points, and reliability priorities. Act as a trusted engineering partner to the desk, ensuring systems are stable, performant, and aligned with business needs. Support live trading environments, including incident response, root cause analysis, and post mortem leadership. Work as a core member of the software engineering team, participating in daily standups and design discussions. Contribute directly to the codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation. Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self healing workflows, and resilience engineering. Use enterprise authorized AI capabilities within the work environment to accelerate major incident triage, troubleshooting, and post incident analysis, validating outputs and handling operational data according to sensitivity and security requirements. Lead reuse first adoption of AI assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls. Drive improvements in latency, throughput, and stability across high volume trading applications. Build and maintain tooling for monitoring, alerting, and distributed tracing across global environments. Operate within a globally distributed engineering and trading organization, collaborating with teams in EMEA, US, and APAC. Partner with infrastructure, networking, cloud engineering, and cybersecurity teams to ensure end to end reliability. Required Qualifications, Capabilities, and Skills: Strong hands on experience in front office trading environments or similarly high pressure, low latency domains. Proficiency with SRE tooling and techniques, including FIX messaging, Kafka, Grafana, Splunk, ITRS Geneos, Dynatrace, InfluxDB, MQ (IBM MQ or similar), Oracle DB. Demonstrated experience using enterprise authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity. Ability to evaluate AI assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations. Deep knowledge of reliability engineering principles: SLIs/SLOs, real time telemetry, disaster recovery planning, capacity planning, and performance tuning. Experience designing and implementing observability frameworks for mission critical systems. Proven ability to lead incident response and drive long term remediation. Solid programming skills in Python, Java, or Kotlin, with the ability to contribute production grade code. Experience with microservices, distributed systems, and event driven architectures. Strong understanding of CI/CD pipelines, automated testing, and deployment strategies. Comfortable interacting directly with traders and senior stakeholders. Excellent communication skills, especially when translating technical issues into business impact. Ability to operate calmly and decisively in high pressure situations. Strong leadership presence with a collaborative mindset.
21/07/2026
Full time
Our trading technology stack is undergoing a multi year convergence and modernization journey. You will play a pivotal role in shaping our next generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast paced front office environments, enjoys direct interaction with traders, and wants to influence the reliability culture of a major global trading organization. As a Lead Site Reliability Engineer at JPMorgan Chase within J.P. Morgan Asset Management's Trading Technology group, you will be embedded directly within the software engineering team that builds and supports our front office trading platforms. Job Responsibilities: Engage daily with traders across asset classes (Equities, Fixed Income, FX) to understand workflows, pain points, and reliability priorities. Act as a trusted engineering partner to the desk, ensuring systems are stable, performant, and aligned with business needs. Support live trading environments, including incident response, root cause analysis, and post mortem leadership. Work as a core member of the software engineering team, participating in daily standups and design discussions. Contribute directly to the codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation. Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self healing workflows, and resilience engineering. Use enterprise authorized AI capabilities within the work environment to accelerate major incident triage, troubleshooting, and post incident analysis, validating outputs and handling operational data according to sensitivity and security requirements. Lead reuse first adoption of AI assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls. Drive improvements in latency, throughput, and stability across high volume trading applications. Build and maintain tooling for monitoring, alerting, and distributed tracing across global environments. Operate within a globally distributed engineering and trading organization, collaborating with teams in EMEA, US, and APAC. Partner with infrastructure, networking, cloud engineering, and cybersecurity teams to ensure end to end reliability. Required Qualifications, Capabilities, and Skills: Strong hands on experience in front office trading environments or similarly high pressure, low latency domains. Proficiency with SRE tooling and techniques, including FIX messaging, Kafka, Grafana, Splunk, ITRS Geneos, Dynatrace, InfluxDB, MQ (IBM MQ or similar), Oracle DB. Demonstrated experience using enterprise authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity. Ability to evaluate AI assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations. Deep knowledge of reliability engineering principles: SLIs/SLOs, real time telemetry, disaster recovery planning, capacity planning, and performance tuning. Experience designing and implementing observability frameworks for mission critical systems. Proven ability to lead incident response and drive long term remediation. Solid programming skills in Python, Java, or Kotlin, with the ability to contribute production grade code. Experience with microservices, distributed systems, and event driven architectures. Strong understanding of CI/CD pipelines, automated testing, and deployment strategies. Comfortable interacting directly with traders and senior stakeholders. Excellent communication skills, especially when translating technical issues into business impact. Ability to operate calmly and decisively in high pressure situations. Strong leadership presence with a collaborative mindset.
IBM
Senior Software Engineer - Confluent Cloud Platform
IBM
At IBM Software, we transform client challenges into solutions. Building the world's leading AI-powered, cloud-native products that shape the future of business and society. Our legacy of innovation creates endless opportunities for IBMers to learn, grow, and make an impact on a global scale. Working in Software means joining a team fueled by curiosity and collaboration. You'll work with diverse technologies, partners, and industries to design, develop, and deliver solutions that power digital transformation. With a culture that values innovation, growth, and continuous learning, IBM Software places you at the heart of IBM's product and technology landscape. Here, you'll have the tools and opportunities to advance your career while creating software that changes the world. With Confluent, data doesn't sit still. We put information in motion, streaming in near real time so organizations can react faster, build smarter, and deliver experiences as dynamic as the world around them. Overview Role and responsibilities at a glance for a Confluent Cloud Infrastructure Software Engineer within the Confluent Cloud Platform and PaaS product. Responsibilities Design, implement and maintain Golang infrastructure services (typically implemented as Kubernetes operators) to deliver the Confluent cloud foundations to the wider engineering organization. Work with Terraform, Datadog, Prometheus; maintain a strong command of Linux, public cloud and networking; Golang software engineering will be the primary focus. Collaborate with the Confluent engineering team to build the PaaS product. Ensure availability, performance, monitoring, emergency response, and capacity planning of the Confluent cloud. Contribute to architecture and operational design points to optimize large-scale data systems (thousands of instances) for security and reliability. Qualifications What You Will Bring: BS, MS, or PhD in computer science or a related field, or equivalent work experience. Relevant cloud infrastructure/cloud networking experience. Strong fundamentals in distributed systems design and development. Experience building and operating large-scale systems. Solid understanding of basic systems operations (disk, network, operating systems, etc.). A self-starter with ability to work effectively in teams. Proficiency in Go, Python, C++, or other statically typed languages. Experience/knowledge with public clouds (AWS, Azure or GCP). What Gives You an Edge Experience with Golang is advantageous. Experience using Apache Kafka is a big plus. About Business Unit IBM Software infuses core business operations with intelligence-from machine learning to generative AI-to help make organizations more responsive, productive, and resilient. IBM Software helps clients put AI into action now to create real value with trust, speed, and confidence across digital labor, IT automation, application modernization, security, and sustainability. Critical to this is the ability to make use of all data, because AI is only as good as the data that fuels it. In most organizations data is spread across multiple clouds, on premises, in private datacenters, and at the edge. IBM's AI and data platform scales and accelerates the impact of AI with trusted data, and provides leading capabilities to train, tune and deploy AI across business. IBM's hybrid cloud platform is one of the most comprehensive and consistent approaches to development, security, and operations across hybrid environments-a flexible foundation for leveraging data, wherever it resides, to extend AI deep into a business. Your In a world where technology never stands still, we understand that dedication to our clients' success, innovation that matters, and trust and personal responsibility in all our relationships, lives in what we do as IBMers as we strive to be the catalyst that makes the world work better. Being an IBMer means you'll be able to learn and develop yourself and your career, you'll be encouraged to be courageous and experiment every day, all whilst having continuous trust and support in an environment where everyone can thrive regardless of background. Our IBMers are growth minded, always staying curious, open to feedback and learning new information and skills to constantly transform themselves and our company. They are trusted to provide ongoing feedback to help other IBMers grow, as well as collaborate with colleagues with a team-focused approach to include different perspectives to drive exceptional outcomes for our customers. The courage our IBMers have to make critical decisions every day is essential to IBM becoming the catalyst for progress, always embracing challenges with available resources, a can-do attitude, and an outcome-focused approach. Are you ready to be an IBMer? About IBM IBM's greatest invention is the IBMer. We believe that through the application of intelligence, reason and science, we can improve business, society and the human condition, bringing the power of an open hybrid cloud and AI strategy to life for our clients and partners around the world. Restlessly reinventing since 1911, we are not only one of the largest corporate organizations in the world, we're also one of the biggest technology and consulting employers, with many of the Fortune 500 companies relying on the IBM Cloud to run their business. At IBM, we pride ourselves on being an early adopter of artificial intelligence, quantum computing and blockchain. Now it's time for you to join us on our journey to being a responsible technology innovator and a force for good in the world. IBM is proud to be an equal-opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, gender, gender identity or expression, sexual orientation, national origin, genetics, pregnancy, disability, neurodivergence, age, or other characteristics protected by the applicable law. IBM is also committed to compliance with all fair employment practices regarding citizenship and immigration status. Other Relevant Job Details For additional information about location requirements, please discuss with the recruiter following submission of your application.
20/07/2026
Full time
At IBM Software, we transform client challenges into solutions. Building the world's leading AI-powered, cloud-native products that shape the future of business and society. Our legacy of innovation creates endless opportunities for IBMers to learn, grow, and make an impact on a global scale. Working in Software means joining a team fueled by curiosity and collaboration. You'll work with diverse technologies, partners, and industries to design, develop, and deliver solutions that power digital transformation. With a culture that values innovation, growth, and continuous learning, IBM Software places you at the heart of IBM's product and technology landscape. Here, you'll have the tools and opportunities to advance your career while creating software that changes the world. With Confluent, data doesn't sit still. We put information in motion, streaming in near real time so organizations can react faster, build smarter, and deliver experiences as dynamic as the world around them. Overview Role and responsibilities at a glance for a Confluent Cloud Infrastructure Software Engineer within the Confluent Cloud Platform and PaaS product. Responsibilities Design, implement and maintain Golang infrastructure services (typically implemented as Kubernetes operators) to deliver the Confluent cloud foundations to the wider engineering organization. Work with Terraform, Datadog, Prometheus; maintain a strong command of Linux, public cloud and networking; Golang software engineering will be the primary focus. Collaborate with the Confluent engineering team to build the PaaS product. Ensure availability, performance, monitoring, emergency response, and capacity planning of the Confluent cloud. Contribute to architecture and operational design points to optimize large-scale data systems (thousands of instances) for security and reliability. Qualifications What You Will Bring: BS, MS, or PhD in computer science or a related field, or equivalent work experience. Relevant cloud infrastructure/cloud networking experience. Strong fundamentals in distributed systems design and development. Experience building and operating large-scale systems. Solid understanding of basic systems operations (disk, network, operating systems, etc.). A self-starter with ability to work effectively in teams. Proficiency in Go, Python, C++, or other statically typed languages. Experience/knowledge with public clouds (AWS, Azure or GCP). What Gives You an Edge Experience with Golang is advantageous. Experience using Apache Kafka is a big plus. About Business Unit IBM Software infuses core business operations with intelligence-from machine learning to generative AI-to help make organizations more responsive, productive, and resilient. IBM Software helps clients put AI into action now to create real value with trust, speed, and confidence across digital labor, IT automation, application modernization, security, and sustainability. Critical to this is the ability to make use of all data, because AI is only as good as the data that fuels it. In most organizations data is spread across multiple clouds, on premises, in private datacenters, and at the edge. IBM's AI and data platform scales and accelerates the impact of AI with trusted data, and provides leading capabilities to train, tune and deploy AI across business. IBM's hybrid cloud platform is one of the most comprehensive and consistent approaches to development, security, and operations across hybrid environments-a flexible foundation for leveraging data, wherever it resides, to extend AI deep into a business. Your In a world where technology never stands still, we understand that dedication to our clients' success, innovation that matters, and trust and personal responsibility in all our relationships, lives in what we do as IBMers as we strive to be the catalyst that makes the world work better. Being an IBMer means you'll be able to learn and develop yourself and your career, you'll be encouraged to be courageous and experiment every day, all whilst having continuous trust and support in an environment where everyone can thrive regardless of background. Our IBMers are growth minded, always staying curious, open to feedback and learning new information and skills to constantly transform themselves and our company. They are trusted to provide ongoing feedback to help other IBMers grow, as well as collaborate with colleagues with a team-focused approach to include different perspectives to drive exceptional outcomes for our customers. The courage our IBMers have to make critical decisions every day is essential to IBM becoming the catalyst for progress, always embracing challenges with available resources, a can-do attitude, and an outcome-focused approach. Are you ready to be an IBMer? About IBM IBM's greatest invention is the IBMer. We believe that through the application of intelligence, reason and science, we can improve business, society and the human condition, bringing the power of an open hybrid cloud and AI strategy to life for our clients and partners around the world. Restlessly reinventing since 1911, we are not only one of the largest corporate organizations in the world, we're also one of the biggest technology and consulting employers, with many of the Fortune 500 companies relying on the IBM Cloud to run their business. At IBM, we pride ourselves on being an early adopter of artificial intelligence, quantum computing and blockchain. Now it's time for you to join us on our journey to being a responsible technology innovator and a force for good in the world. IBM is proud to be an equal-opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, gender, gender identity or expression, sexual orientation, national origin, genetics, pregnancy, disability, neurodivergence, age, or other characteristics protected by the applicable law. IBM is also committed to compliance with all fair employment practices regarding citizenship and immigration status. Other Relevant Job Details For additional information about location requirements, please discuss with the recruiter following submission of your application.
Senior Manager, Center of Expertise
Veeam Software
Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. As the market leader in both data resilience and data security posture management, Veeam is built for the convergence of identity, data, security, and AI risk. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide, who trust Veeam to keep their businesses running. Join us as we go fearlessly forward together, growing, learning, and making a real impact for some of the world's biggest brands. About the Role The Senior Manager, Center of Expertise will lead a team of elite product and domain experts who drive measurable outcomes on strategic customer engagements. This role is built for a technically fluent, decisive leader who thrives at the intersection of modern data protection, cloud infrastructure, and AI driven operations-someone who can translate complex architectures into executive level narratives and who genuinely cares about the people they lead and the customers they serve. You will lead a team of Domain Engineering Specialists (DES) who serve as subject matter experts across Veeam Data Platform (VDP), Veeam Data Cloud (VDC), Vault, and Kasten, as well as the broader cloud native and AI/ML ecosystem surrounding them. This is a team that sets the standard-not just within Veeam, but across the industry-and the right leader will take pride in keeping it that way. Your team partners with customer facing teams to lead complex, large scale technical onboardings involving intricate multi environment architectures, data modeling, telemetry analysis, and risk discussions with C suite stakeholders (CISO/CIO/CTO). You will align customer posture against the Veeam Data Resilience Maturity Model (DRMM) and industry frameworks (NIST CSF, CIS Controls, Zero Trust Architecture) to surface optimization opportunities and drive expansion. Success requires deep cross functional alignment, the ability to manage competing priorities at scale, and a bias toward outcomes over activity. What You'll Do Lead and develop a team of 6-10 Domain Engineering Specialists, investing genuinely in their growth through technical coaching, structured career development, and a culture where people feel challenged, valued, and proud of the work they do. Define and execute the team's technical enablement roadmap, with a strong focus on emerging capabilities in AI/ML, cloud native architectures, and cyber resilience. Own the team's involvement in complex, large scale customer onboardings-multi environment migrations, hybrid and multi cloud transformations, and enterprise wide resilience deployments that require deep architectural engagement and precise coordination. Coach specialists to deepen expertise in ransomware recovery, immutable storage, multi cloud data protection, and Kubernetes native backup-and create clear paths for them to grow into recognized experts inside and outside of Veeam. Build and scale a content and resource hub that extends the team's reach beyond 1:1 customer engagements-including technical playbooks, reference architectures, self service assessment tools, and evergreen enablement content. Develop and run high impact customer engagement programs-webinars, roundtables, virtual workshops, executive briefings, and community forums-that position Veeam's expertise in front of broader audiences and create scalable touchpoints across the customer lifecycle. Establish and maintain the team's position as a credible technical authority in conversations with customer engineering, architecture, and executive stakeholders. Drive process improvements and playbooks that scale impact without scaling headcount linearly. Monitor operational KPIs and use data to identify coaching opportunities, capacity gaps, and emerging risk patterns across the customer portfolio. Influence product and engineering roadmaps by synthesizing field level insights into structured feedback loops. Build and maintain cross functional relationships with Sales, Product, Engineering, and Customer Success to ensure coordinated customer outcomes. Serve as an executive level escalation point for complex technical and strategic customer situations. Model and reinforce a high performance culture anchored in intellectual curiosity, ownership, and genuine care for customers and teammates alike. And other responsibilities as needed as the business evolves and grows. What You'll Bring 8+ years of experience in technical customer success, solutions engineering, or related customer facing technical roles-ideally within enterprise SaaS, cloud infrastructure, or data management. 4+ years of people management experience, with a track record of building and scaling high performing technical teams where culture is a competitive advantage, not an afterthought. Deep expertise in data protection and cyber resilience-including ransomware recovery architectures, immutability strategies, air gapped backup, and recovery time/point objectives at enterprise scale. Hands on familiarity with AI and machine learning concepts, including how organizations are operationalizing AI workloads and the data infrastructure that supports them (data pipelines, model training environments, vector databases, LLMOps). Strong working knowledge of major cloud platforms (AWS, Azure, GCP) and cloud native patterns-Kubernetes, containers, serverless, IaC (Terraform/Pulumi), and multi cloud data governance. Fluency in modern infrastructure concepts: hybrid cloud architectures, software defined storage, disaster recovery as code, and observability stacks. Ability to engage credibly in security framework discussions-Zero Trust, NIST CSF, CIS Controls, SOC 2, and regulatory environments (GDPR, HIPAA, FedRAMP). Experience leading or owning large scale, complex customer onboardings with multiple stakeholders, integrated environments, and significant technical and organizational complexity. Proven ability to build scalable content programs-technical resource hubs, reference architectures, webinars, workshops-that extend expert knowledge to audiences well beyond direct engagement. Demonstrated ability to lead teams through ambiguity in fast moving, high growth environments. Strong executive communication skills-able to distill complex technical risk into business impact narratives for CISO, CIO, and board level audiences. Experience using data and telemetry to drive coaching decisions and team performance improvements. Adaptable, high agency operator who sets direction without waiting for perfect information-and brings others along with them. What you'll get 25 paid vacation days, plus 4 extra global VeeaMe Days for self care and 24 paid volunteer hours annually through Veeam Cares. Private medical, dental, and vision insurance with dependent enrolment. Life insurance with enhanced coverage and global 24/7 protection. Income protection after 26 weeks, covering a portion of salary. Defined contribution pension plan with employer match. Worldwide travel insurance for business and leisure, with option to enroll dependents. Employee Assistance Program with therapy, legal, and financial support, plus online GP services and wellbeing programs. Opportunities to learn and grow through on demand libraries (LinkedIn Learning, O'Reilly), mentoring, workshops and learning events like our annual Global Day of Learning. Veeam Software is an equal opportunity employer Veeam Software is an equal opportunity employer and does not tolerate discrimination in any form on the basis of race, color, religion, gender, age, national origin, citizenship, disability, veteran status or any other classification protected by federal, state or local law. All your information will be kept confidential.
19/07/2026
Full time
Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. As the market leader in both data resilience and data security posture management, Veeam is built for the convergence of identity, data, security, and AI risk. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide, who trust Veeam to keep their businesses running. Join us as we go fearlessly forward together, growing, learning, and making a real impact for some of the world's biggest brands. About the Role The Senior Manager, Center of Expertise will lead a team of elite product and domain experts who drive measurable outcomes on strategic customer engagements. This role is built for a technically fluent, decisive leader who thrives at the intersection of modern data protection, cloud infrastructure, and AI driven operations-someone who can translate complex architectures into executive level narratives and who genuinely cares about the people they lead and the customers they serve. You will lead a team of Domain Engineering Specialists (DES) who serve as subject matter experts across Veeam Data Platform (VDP), Veeam Data Cloud (VDC), Vault, and Kasten, as well as the broader cloud native and AI/ML ecosystem surrounding them. This is a team that sets the standard-not just within Veeam, but across the industry-and the right leader will take pride in keeping it that way. Your team partners with customer facing teams to lead complex, large scale technical onboardings involving intricate multi environment architectures, data modeling, telemetry analysis, and risk discussions with C suite stakeholders (CISO/CIO/CTO). You will align customer posture against the Veeam Data Resilience Maturity Model (DRMM) and industry frameworks (NIST CSF, CIS Controls, Zero Trust Architecture) to surface optimization opportunities and drive expansion. Success requires deep cross functional alignment, the ability to manage competing priorities at scale, and a bias toward outcomes over activity. What You'll Do Lead and develop a team of 6-10 Domain Engineering Specialists, investing genuinely in their growth through technical coaching, structured career development, and a culture where people feel challenged, valued, and proud of the work they do. Define and execute the team's technical enablement roadmap, with a strong focus on emerging capabilities in AI/ML, cloud native architectures, and cyber resilience. Own the team's involvement in complex, large scale customer onboardings-multi environment migrations, hybrid and multi cloud transformations, and enterprise wide resilience deployments that require deep architectural engagement and precise coordination. Coach specialists to deepen expertise in ransomware recovery, immutable storage, multi cloud data protection, and Kubernetes native backup-and create clear paths for them to grow into recognized experts inside and outside of Veeam. Build and scale a content and resource hub that extends the team's reach beyond 1:1 customer engagements-including technical playbooks, reference architectures, self service assessment tools, and evergreen enablement content. Develop and run high impact customer engagement programs-webinars, roundtables, virtual workshops, executive briefings, and community forums-that position Veeam's expertise in front of broader audiences and create scalable touchpoints across the customer lifecycle. Establish and maintain the team's position as a credible technical authority in conversations with customer engineering, architecture, and executive stakeholders. Drive process improvements and playbooks that scale impact without scaling headcount linearly. Monitor operational KPIs and use data to identify coaching opportunities, capacity gaps, and emerging risk patterns across the customer portfolio. Influence product and engineering roadmaps by synthesizing field level insights into structured feedback loops. Build and maintain cross functional relationships with Sales, Product, Engineering, and Customer Success to ensure coordinated customer outcomes. Serve as an executive level escalation point for complex technical and strategic customer situations. Model and reinforce a high performance culture anchored in intellectual curiosity, ownership, and genuine care for customers and teammates alike. And other responsibilities as needed as the business evolves and grows. What You'll Bring 8+ years of experience in technical customer success, solutions engineering, or related customer facing technical roles-ideally within enterprise SaaS, cloud infrastructure, or data management. 4+ years of people management experience, with a track record of building and scaling high performing technical teams where culture is a competitive advantage, not an afterthought. Deep expertise in data protection and cyber resilience-including ransomware recovery architectures, immutability strategies, air gapped backup, and recovery time/point objectives at enterprise scale. Hands on familiarity with AI and machine learning concepts, including how organizations are operationalizing AI workloads and the data infrastructure that supports them (data pipelines, model training environments, vector databases, LLMOps). Strong working knowledge of major cloud platforms (AWS, Azure, GCP) and cloud native patterns-Kubernetes, containers, serverless, IaC (Terraform/Pulumi), and multi cloud data governance. Fluency in modern infrastructure concepts: hybrid cloud architectures, software defined storage, disaster recovery as code, and observability stacks. Ability to engage credibly in security framework discussions-Zero Trust, NIST CSF, CIS Controls, SOC 2, and regulatory environments (GDPR, HIPAA, FedRAMP). Experience leading or owning large scale, complex customer onboardings with multiple stakeholders, integrated environments, and significant technical and organizational complexity. Proven ability to build scalable content programs-technical resource hubs, reference architectures, webinars, workshops-that extend expert knowledge to audiences well beyond direct engagement. Demonstrated ability to lead teams through ambiguity in fast moving, high growth environments. Strong executive communication skills-able to distill complex technical risk into business impact narratives for CISO, CIO, and board level audiences. Experience using data and telemetry to drive coaching decisions and team performance improvements. Adaptable, high agency operator who sets direction without waiting for perfect information-and brings others along with them. What you'll get 25 paid vacation days, plus 4 extra global VeeaMe Days for self care and 24 paid volunteer hours annually through Veeam Cares. Private medical, dental, and vision insurance with dependent enrolment. Life insurance with enhanced coverage and global 24/7 protection. Income protection after 26 weeks, covering a portion of salary. Defined contribution pension plan with employer match. Worldwide travel insurance for business and leisure, with option to enroll dependents. Employee Assistance Program with therapy, legal, and financial support, plus online GP services and wellbeing programs. Opportunities to learn and grow through on demand libraries (LinkedIn Learning, O'Reilly), mentoring, workshops and learning events like our annual Global Day of Learning. Veeam Software is an equal opportunity employer Veeam Software is an equal opportunity employer and does not tolerate discrimination in any form on the basis of race, color, religion, gender, age, national origin, citizenship, disability, veteran status or any other classification protected by federal, state or local law. All your information will be kept confidential.
Member of Technical Staff - Engineering Lead, Compute Platform
Reflection
Our Mission Reflection is a research lab making intelligence open and accessible for everyone to use, customize, and build on. We build open models that let anyone control their intelligence and help shape the future of AI. Our mission: make intelligence open and accessible to all. ABOUT THE ROLE Reflection's Compute Platform team keeps our compute layer healthy and highly available. We run a Kubernetes-based platform distributed across multiple neo-clouds, where multi-cloud scheduling, node health, and performance debugging at scale present genuinely hard systems problems. As Compute Platform Lead, you'll provide front-line leadership of the team that builds and operates this layer. You'll build, mentor, and grow a team of strong systems engineers, guide the technical and architectural decisions across multi-cloud scheduling, cluster management, and next-generation GPU deployments, and work closely with our training teams to co-design fault tolerance, node health checks, and remediation. You'll stay close enough to the systems to make targeted contributions as an individual contributor and to maintain a deep understanding of the compute fleet our largest training runs depend on. Managing vendors - and the important deals that come with them - is a core part of the job. WHAT YOU'LL DO Build, mentor, and grow a high-performing team of systems engineers. Coach and support your reports in understanding, and pursuing, their professional growth. Provide front-line leadership of engineering efforts to keep the compute fleet reliable and highly available - multi-cloud scheduling, cluster management, and the path to next-generation GPUs and increasingly larger cluster sizes. Stay hands-on: become familiar with the team's technical stack enough to make targeted contributions as an individual contributor. Manage day-to-day execution: prioritize the team's work and manage projects in a highly dynamic, fast-paced environment. Guide technical and architectural decisions, emphasizing scalability, robustness, and reliability - automatic remediation, topology-aware scheduling, capacity planning, rapid hardware debugging, and cluster-wide monitoring and performance benchmarking. Work closely with our training teams to co-design fault tolerance, node health checks, and remediation, and manage the vendor relationships and important deals the compute fleet depends on. Prepare the fleet for what's next: next-generation GPUs and larger clusters, and - longer term - multi-cloud storage, petabyte-scale data replication, and GPU-to-GPU network performance. Raise the bar for technical judgment, prioritization, communication, and execution in a fast-moving environment. WHAT WE'RE LOOKING FOR Experience building, mentoring, and growing systems or infrastructure teams while staying technically hands-on. (Comfortable growing into leading a team of 10 quickly if you haven't managed at that scale before.) Deep systems-level engineering experience with a focus on cluster-wide behavior and maintenance. Strong coding ability and the credibility to earn the technical trust of a strong team. Depth in at least one of orchestration, storage, or GPU hardware - with the ability to learn the rest. Deep GPU knowledge beyond standard Kubernetes (e.g., NCCL) is a plus, not a prerequisite. Alignment with a Kubernetes-first architecture. Cloud storage expertise - managing high-performance data products (like VAST) across multiple data centers and handling datasets and checkpointing at scale - is a plus. Experience managing vendors, including negotiating and operating important deals. Ability to guide strategy and drive execution across a multi-cloud, large-fleet environment, and to partner effectively with research and training teams. What We Offer: We believe that to make intelligence open and accessible to all, you need to start at the foundation. Joining Reflection means building from the ground up as part of a talent-dense team. You will help define our future as a company, and help define the future of open foundational models. We want you to do the most impactful work of your career with the confidence that you and the people you care about most are supported. Top-tier compensation: Salary and equity structured to recognize and retain our talent globally. Stock options: Everyone who joins and contributes to Reflection's success gets to share in the upside through stock options. Health & wellness: Comprehensive medical, dental, vision, and life, with an annual wellness allowance. Meals: Lunch and dinner are provided in the office daily. Life & family: 22 weeks paid parental leave for all new birthing and non-birthing parents, including adoptive and surrogate journeys. Vacation days: Unlimited paid time off in the U.S. and 30 days in the U.K. Sponsorship support: We sponsor visas to help exceptional talent join our team and support long-term immigration pathways where applicable. Team building: We have regular off-sites, happy hours, and team celebrations. Export Control Notice: This position may require access to technology or source code subject to the U.S. Export Administration Regulations. Any offer of employment for this role may be conditioned on the Company's ability to provide the candidate with access to such technology or source code in compliance with applicable U.S. export control laws, which may require the Company to seek government authorization.
19/07/2026
Full time
Our Mission Reflection is a research lab making intelligence open and accessible for everyone to use, customize, and build on. We build open models that let anyone control their intelligence and help shape the future of AI. Our mission: make intelligence open and accessible to all. ABOUT THE ROLE Reflection's Compute Platform team keeps our compute layer healthy and highly available. We run a Kubernetes-based platform distributed across multiple neo-clouds, where multi-cloud scheduling, node health, and performance debugging at scale present genuinely hard systems problems. As Compute Platform Lead, you'll provide front-line leadership of the team that builds and operates this layer. You'll build, mentor, and grow a team of strong systems engineers, guide the technical and architectural decisions across multi-cloud scheduling, cluster management, and next-generation GPU deployments, and work closely with our training teams to co-design fault tolerance, node health checks, and remediation. You'll stay close enough to the systems to make targeted contributions as an individual contributor and to maintain a deep understanding of the compute fleet our largest training runs depend on. Managing vendors - and the important deals that come with them - is a core part of the job. WHAT YOU'LL DO Build, mentor, and grow a high-performing team of systems engineers. Coach and support your reports in understanding, and pursuing, their professional growth. Provide front-line leadership of engineering efforts to keep the compute fleet reliable and highly available - multi-cloud scheduling, cluster management, and the path to next-generation GPUs and increasingly larger cluster sizes. Stay hands-on: become familiar with the team's technical stack enough to make targeted contributions as an individual contributor. Manage day-to-day execution: prioritize the team's work and manage projects in a highly dynamic, fast-paced environment. Guide technical and architectural decisions, emphasizing scalability, robustness, and reliability - automatic remediation, topology-aware scheduling, capacity planning, rapid hardware debugging, and cluster-wide monitoring and performance benchmarking. Work closely with our training teams to co-design fault tolerance, node health checks, and remediation, and manage the vendor relationships and important deals the compute fleet depends on. Prepare the fleet for what's next: next-generation GPUs and larger clusters, and - longer term - multi-cloud storage, petabyte-scale data replication, and GPU-to-GPU network performance. Raise the bar for technical judgment, prioritization, communication, and execution in a fast-moving environment. WHAT WE'RE LOOKING FOR Experience building, mentoring, and growing systems or infrastructure teams while staying technically hands-on. (Comfortable growing into leading a team of 10 quickly if you haven't managed at that scale before.) Deep systems-level engineering experience with a focus on cluster-wide behavior and maintenance. Strong coding ability and the credibility to earn the technical trust of a strong team. Depth in at least one of orchestration, storage, or GPU hardware - with the ability to learn the rest. Deep GPU knowledge beyond standard Kubernetes (e.g., NCCL) is a plus, not a prerequisite. Alignment with a Kubernetes-first architecture. Cloud storage expertise - managing high-performance data products (like VAST) across multiple data centers and handling datasets and checkpointing at scale - is a plus. Experience managing vendors, including negotiating and operating important deals. Ability to guide strategy and drive execution across a multi-cloud, large-fleet environment, and to partner effectively with research and training teams. What We Offer: We believe that to make intelligence open and accessible to all, you need to start at the foundation. Joining Reflection means building from the ground up as part of a talent-dense team. You will help define our future as a company, and help define the future of open foundational models. We want you to do the most impactful work of your career with the confidence that you and the people you care about most are supported. Top-tier compensation: Salary and equity structured to recognize and retain our talent globally. Stock options: Everyone who joins and contributes to Reflection's success gets to share in the upside through stock options. Health & wellness: Comprehensive medical, dental, vision, and life, with an annual wellness allowance. Meals: Lunch and dinner are provided in the office daily. Life & family: 22 weeks paid parental leave for all new birthing and non-birthing parents, including adoptive and surrogate journeys. Vacation days: Unlimited paid time off in the U.S. and 30 days in the U.K. Sponsorship support: We sponsor visas to help exceptional talent join our team and support long-term immigration pathways where applicable. Team building: We have regular off-sites, happy hours, and team celebrations. Export Control Notice: This position may require access to technology or source code subject to the U.S. Export Administration Regulations. Any offer of employment for this role may be conditioned on the Company's ability to provide the candidate with access to such technology or source code in compliance with applicable U.S. export control laws, which may require the Company to seek government authorization.

Modal Window

  • Home
  • Contact
  • About Us
  • FAQs
  • Terms & Conditions
  • Privacy
  • Employer
  • Post a Job
  • Search Resumes
  • Sign in
  • Job Seeker
  • Find Jobs
  • Create Resume
  • Sign in
  • IT blog
  • Facebook
  • Twitter
  • LinkedIn
  • Youtube
© 2008-2026 IT Job Board