My leading Tech client are looking for a talented and motivated individual to ensure the resilience, performance, and cost-effectiveness of their Azure-based data platform. This role is essential to their data ecosystem, combining platform reliability, incident response, SLA management, cost optimization (FinOps), and deployment oversight. You will be the single point of contact for operational issues, driving rapid resolution during outages, leading communications with stakeholders, and shaping the processes that keeps their platform running smoothly and efficiently. This is a newly created role in a growing business. A brilliant opportunity! The following skills/experience is required: Proven operational leadership for large-scale data platforms. Expertise in incident management, SLA enforcement, and stakeholder communication. Hands-on experience with Azure Synapse, Databricks, ADF, Power BI. Familiarity with CI/CD and automation. Strong FinOps mindset and cost management experience. Knowledge of monitoring and observability frameworks. Salary: Up to £120,000 + bonus + package Level: Manager Location: London (good work from home options available) If you are interested in this Data Ops Manager position and meet the above requirements please apply immediately.
27/07/2026
Full time
My leading Tech client are looking for a talented and motivated individual to ensure the resilience, performance, and cost-effectiveness of their Azure-based data platform. This role is essential to their data ecosystem, combining platform reliability, incident response, SLA management, cost optimization (FinOps), and deployment oversight. You will be the single point of contact for operational issues, driving rapid resolution during outages, leading communications with stakeholders, and shaping the processes that keeps their platform running smoothly and efficiently. This is a newly created role in a growing business. A brilliant opportunity! The following skills/experience is required: Proven operational leadership for large-scale data platforms. Expertise in incident management, SLA enforcement, and stakeholder communication. Hands-on experience with Azure Synapse, Databricks, ADF, Power BI. Familiarity with CI/CD and automation. Strong FinOps mindset and cost management experience. Knowledge of monitoring and observability frameworks. Salary: Up to £120,000 + bonus + package Level: Manager Location: London (good work from home options available) If you are interested in this Data Ops Manager position and meet the above requirements please apply immediately.
Overview The Data Engineering Manager is responsible for defining and executing a modern, scalable, and governed data engineering strategy to support trading, analytics, and AI/ML use cases across the organization. This role sits within the core technology function, reporting to the Head of Technology, and works with platform, infrastructure, and security teams to build reliable, governed, high-performance data systems that serve the entire business. The focus is on delivering high-quality, reliable, and reusable data products through event-driven architectures and high-throughput data pipelines. The role combines hands-on data engineering expertise with full ownership of the underlying data platform and its integration within the broader infrastructure landscape, ensuring alignment with enterprise architecture, security, and operational standards. Working with Trading, Risk, Finance, and Technology teams, this role ensures that complex data is transformed into trusted, accessible, and performant datasets aligned with governance frameworks and the firm's semantic/knowledge-graph vision. Responsibilities Define and execute the end-to-end data engineering strategy (ingest govern serve observe), aligned with enterprise architecture, platform standards, and data governance. Own the design, build, and operation of the data platform on Azure and Snowflake, including integration with infrastructure components, ensuring scalability, security, and reliability. Design and implement event-driven architectures and high-throughput data pipelines to support real-time and batch data processing. Develop and industrialise high-performance data pipelines across key domains such as market/curve data, ETRM/CTRM systems, and finance/settlement, ensuring SLOs, lineage, and DR/BCP requirements are met. Define and enforce engineering standards across data pipelines, data models, APIs, and serving layers (SQL/API/Graph) to ensure consistency, reuse, and scalability. Embed data governance by design, including data contracts, data quality rules, access controls, masking, retention, and compliance requirements. Drive performance optimisation and cloud cost optimisation through efficient architecture and FinOps practices. Own the reliability and operational excellence of the data platform, including observability (metrics, logs, traces), incident management, and continuous improvement of critical data flows. Lead, coach, and scale a global team of data engineers, ensuring strong technical standards, delivery excellence, and a high-performing engineering culture. Act as the key interface between data engineering and the broader technology ecosystem (platform, infrastructure, security), ensuring seamless integration and alignment with enterprise standards. Partner with Business to prioritise initiatives and deliver impactful, production-grade data solutions. Profile Bachelor's degree or higher in Computer Science, Engineering, Applied Mathematics, or a related field. 7-10 years of experience in data engineering or data platform roles, including team leadership in complex, real-time environments. Strong hands-on experience with Azure as well as Snowflake, including performance, security, and cost optimisation. Solid experience with modern data architectures and tools (i.e. Terraform, Dagster). Proven ability to build and operate scalable, high-performance data pipelines with strong focus on reliability and observability. Strong understanding of platform and infrastructure concepts. Experience with data governance practices and enabling data for AI/ML use cases. Strong stakeholder management and communication skills, with the ability to work across technical and business teams. Ability to manage and prioritize multiple tasks in a fast-paced, deadline-driven environment. Energy commodity trading experience is a real advantage. Other cloud platform knowledge is a plus. Fluent in English (spoken and written). If you think the open position you see is right for you, we encourage you to apply. Our people make all the difference in our success.
27/07/2026
Full time
Overview The Data Engineering Manager is responsible for defining and executing a modern, scalable, and governed data engineering strategy to support trading, analytics, and AI/ML use cases across the organization. This role sits within the core technology function, reporting to the Head of Technology, and works with platform, infrastructure, and security teams to build reliable, governed, high-performance data systems that serve the entire business. The focus is on delivering high-quality, reliable, and reusable data products through event-driven architectures and high-throughput data pipelines. The role combines hands-on data engineering expertise with full ownership of the underlying data platform and its integration within the broader infrastructure landscape, ensuring alignment with enterprise architecture, security, and operational standards. Working with Trading, Risk, Finance, and Technology teams, this role ensures that complex data is transformed into trusted, accessible, and performant datasets aligned with governance frameworks and the firm's semantic/knowledge-graph vision. Responsibilities Define and execute the end-to-end data engineering strategy (ingest govern serve observe), aligned with enterprise architecture, platform standards, and data governance. Own the design, build, and operation of the data platform on Azure and Snowflake, including integration with infrastructure components, ensuring scalability, security, and reliability. Design and implement event-driven architectures and high-throughput data pipelines to support real-time and batch data processing. Develop and industrialise high-performance data pipelines across key domains such as market/curve data, ETRM/CTRM systems, and finance/settlement, ensuring SLOs, lineage, and DR/BCP requirements are met. Define and enforce engineering standards across data pipelines, data models, APIs, and serving layers (SQL/API/Graph) to ensure consistency, reuse, and scalability. Embed data governance by design, including data contracts, data quality rules, access controls, masking, retention, and compliance requirements. Drive performance optimisation and cloud cost optimisation through efficient architecture and FinOps practices. Own the reliability and operational excellence of the data platform, including observability (metrics, logs, traces), incident management, and continuous improvement of critical data flows. Lead, coach, and scale a global team of data engineers, ensuring strong technical standards, delivery excellence, and a high-performing engineering culture. Act as the key interface between data engineering and the broader technology ecosystem (platform, infrastructure, security), ensuring seamless integration and alignment with enterprise standards. Partner with Business to prioritise initiatives and deliver impactful, production-grade data solutions. Profile Bachelor's degree or higher in Computer Science, Engineering, Applied Mathematics, or a related field. 7-10 years of experience in data engineering or data platform roles, including team leadership in complex, real-time environments. Strong hands-on experience with Azure as well as Snowflake, including performance, security, and cost optimisation. Solid experience with modern data architectures and tools (i.e. Terraform, Dagster). Proven ability to build and operate scalable, high-performance data pipelines with strong focus on reliability and observability. Strong understanding of platform and infrastructure concepts. Experience with data governance practices and enabling data for AI/ML use cases. Strong stakeholder management and communication skills, with the ability to work across technical and business teams. Ability to manage and prioritize multiple tasks in a fast-paced, deadline-driven environment. Energy commodity trading experience is a real advantage. Other cloud platform knowledge is a plus. Fluent in English (spoken and written). If you think the open position you see is right for you, we encourage you to apply. Our people make all the difference in our success.
Embrace this pivotal role as an essential member of a high performing team dedicated to reaching new heights in data engineering. Your contributions will be instrumental in shaping the future of one of the world's largest and most influential companies. As a Senior Lead Data Engineer at JPMorganChase within the Behavioral Insights Team, you will turn operational signals and platform data into actionable insights that improve reliability, risk/control health, and delivery efficiency. You will own the reliability and performance of reporting and analytics pipelines, define and govern SLIs/SLOs and error budgets, and automate data products (dashboards, scheduled reporting, and near-real-time views) that serve engineering, risk, and business stakeholders. You will also manage and mentor team members and uphold rigorous data management practices and controls. Job Responsibilities Change and release health - Track deployment frequency, change failure rate, lead time, and rollbacks; correlate changes to incidents/SLO impact; influence safer release practices. Capacity, performance, and scalability - Produce capacity forecasts, headroom and hotspot reporting; partner with engineering to validate scaling policies and performance budgets. FinOps and cost observability - Report spend by service/team/env; track unit economics (e.g., cost per transaction), rightsizing opportunities, commitment utilization, and tag compliance; highlight reliability-cost tradeoffs. Risk and controls compliance - Evidence guardrail adherence and control health (backup/restore posture, DR testing, patch/vulnerability closure, config drift); ensure metrics lineage and audit readiness. Uses enterprise-authorized AI capabilities within the work environment to accelerate data platform and design analysis and technical documentation, validating outputs and handling data according to sensitivity and security requirements. Data risk and controls - Monitor adherence to risk and control guidelines for data access and use. Data quality, reliability, and lineage - Define data contracts for telemetry sources; implement validation, anomaly detection, and reconciliation; document definitions (e.g., formulas, thresholds) and end to end data lineage from raw signals to KPIs and insights. Automation and self service - Deliver automated pipelines for scheduled reporting and near real time dashboards; enable RBAC controlled self service for teams and leadership. Stakeholder cadences and communication - Lead weekly reliability reviews and monthly leadership reviews; maintain action logs to closure; elevate risks early with data driven recommendations. People leadership - oversee workflow, prioritization, and delivery for junior data engineers and visualization researchers; mentor and support upskilling and career development. Applies reuse first, AI assisted practices within delivery and operational routines (e.g., validation automation and access control review support), ensuring traceability/auditability and alignment to resiliency and security expectations. Required qualifications, capabilities, and skills Formal training or certification on data engineering concepts and advanced applied experience in data analytics/BI/ operations analytics. Strong SQL skills (CTEs, window functions, performance aware querying). Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support data engineering workflows with strong validation habits and awareness of data sensitivity. Ability to review and validate AI assisted outputs (e.g., model and design summaries or validation recommendations) before use, escalating when uncertain and following data handling requirements. Hands on experience building dashboards in Tableau/Power BI/Looker (or similar). Experience working with ITSM tools (e.g., ServiceNow or similar) and understanding incident/change/problem concepts. Strong data storytelling skills; ability to translate operational findings into practical improvements. Preferred qualifications Experience working with AWS and core concepts (accounts, regions, IAM, networking, compute/storage, tagging). Python for analytics/automation (e.g., pandas) and building repeatable pipelines. Experience with cloud data platforms (e.g., Snowflake/Redshift/BigQuery) and ELT tooling (e.g., dbt).
27/07/2026
Full time
Embrace this pivotal role as an essential member of a high performing team dedicated to reaching new heights in data engineering. Your contributions will be instrumental in shaping the future of one of the world's largest and most influential companies. As a Senior Lead Data Engineer at JPMorganChase within the Behavioral Insights Team, you will turn operational signals and platform data into actionable insights that improve reliability, risk/control health, and delivery efficiency. You will own the reliability and performance of reporting and analytics pipelines, define and govern SLIs/SLOs and error budgets, and automate data products (dashboards, scheduled reporting, and near-real-time views) that serve engineering, risk, and business stakeholders. You will also manage and mentor team members and uphold rigorous data management practices and controls. Job Responsibilities Change and release health - Track deployment frequency, change failure rate, lead time, and rollbacks; correlate changes to incidents/SLO impact; influence safer release practices. Capacity, performance, and scalability - Produce capacity forecasts, headroom and hotspot reporting; partner with engineering to validate scaling policies and performance budgets. FinOps and cost observability - Report spend by service/team/env; track unit economics (e.g., cost per transaction), rightsizing opportunities, commitment utilization, and tag compliance; highlight reliability-cost tradeoffs. Risk and controls compliance - Evidence guardrail adherence and control health (backup/restore posture, DR testing, patch/vulnerability closure, config drift); ensure metrics lineage and audit readiness. Uses enterprise-authorized AI capabilities within the work environment to accelerate data platform and design analysis and technical documentation, validating outputs and handling data according to sensitivity and security requirements. Data risk and controls - Monitor adherence to risk and control guidelines for data access and use. Data quality, reliability, and lineage - Define data contracts for telemetry sources; implement validation, anomaly detection, and reconciliation; document definitions (e.g., formulas, thresholds) and end to end data lineage from raw signals to KPIs and insights. Automation and self service - Deliver automated pipelines for scheduled reporting and near real time dashboards; enable RBAC controlled self service for teams and leadership. Stakeholder cadences and communication - Lead weekly reliability reviews and monthly leadership reviews; maintain action logs to closure; elevate risks early with data driven recommendations. People leadership - oversee workflow, prioritization, and delivery for junior data engineers and visualization researchers; mentor and support upskilling and career development. Applies reuse first, AI assisted practices within delivery and operational routines (e.g., validation automation and access control review support), ensuring traceability/auditability and alignment to resiliency and security expectations. Required qualifications, capabilities, and skills Formal training or certification on data engineering concepts and advanced applied experience in data analytics/BI/ operations analytics. Strong SQL skills (CTEs, window functions, performance aware querying). Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support data engineering workflows with strong validation habits and awareness of data sensitivity. Ability to review and validate AI assisted outputs (e.g., model and design summaries or validation recommendations) before use, escalating when uncertain and following data handling requirements. Hands on experience building dashboards in Tableau/Power BI/Looker (or similar). Experience working with ITSM tools (e.g., ServiceNow or similar) and understanding incident/change/problem concepts. Strong data storytelling skills; ability to translate operational findings into practical improvements. Preferred qualifications Experience working with AWS and core concepts (accounts, regions, IAM, networking, compute/storage, tagging). Python for analytics/automation (e.g., pandas) and building repeatable pipelines. Experience with cloud data platforms (e.g., Snowflake/Redshift/BigQuery) and ELT tooling (e.g., dbt).
Project overview: The project focuses on guiding large-scale cloud adoption and infrastructure modernization initiatives with an emphasis on structured and secure environments. You will contribute to building robust cloud foundations, enabling scalable architectures, and aligning multi-cloud strategies with long term business goals. Position overview: We are looking for a GCP Solution Architect who can guide complex cloud transformation initiatives with a focus on cloud adoption and infrastructure modernization. You will define architectural standards, align engineering teams with business objectives, and ensure cloud environments are secure, compliant, and scalable, without requiring hands on coding. Technology stack: Google Cloud Platform (GCP), AWS, Azure, GCP Cloud Adoption Framework, IAM, VPC Service Controls, Organization Policies, Databricks, PostgreSQL, dbt, Snowflake, Apigee, CI/CD pipelines, and FinOps practices. Responsibilities Define scalable and resilient cloud architectures aligned with enterprise standards Collaborate with engineering teams to design cloud solutions using structured adoption frameworks Translate business requirements into clear architecture blueprints Develop migration strategies, roadmaps, and cost aware planning models Communicate architectural decisions, risks, and progress to stakeholders in a clear manner Ensure alignment with enterprise governance, compliance, and security standards Identify and mitigate technical and operational risks early in the lifecycle Act as a bridge between business stakeholders, compliance teams, and engineering groups Guide adoption of modern cloud architecture patterns and best practices Support improvements in cloud maturity, scalability, and operational efficiency Requirements Experience working as a Lead or Solution Architect in complex enterprise environments Strong experience designing cloud architectures and infrastructure solutions Knowledge of Google Cloud Platform including landing zones, resource hierarchy, and organizational structure Understanding of cloud security principles including identity management and governance frameworks Experience defining high level architecture blueprints aligned with business strategies Knowledge of data governance practices including data lifecycle management and access control Understanding of regulatory and compliance requirements in financial environments Experience working with multi cloud environments including AWS and Azure Familiarity with tools such as Databricks, PostgreSQL, dbt, Snowflake, and Apigee Strong communication and stakeholder management skillsExperience working within Agile and DevOps environments Nice to have Experience with FinOps practices and cloud cost optimization strategies Exposure to hybrid cloud implementations Experience supporting enterprise scale cloud migration programs Familiarity with observability tools and cloud monitoring solutions
27/07/2026
Full time
Project overview: The project focuses on guiding large-scale cloud adoption and infrastructure modernization initiatives with an emphasis on structured and secure environments. You will contribute to building robust cloud foundations, enabling scalable architectures, and aligning multi-cloud strategies with long term business goals. Position overview: We are looking for a GCP Solution Architect who can guide complex cloud transformation initiatives with a focus on cloud adoption and infrastructure modernization. You will define architectural standards, align engineering teams with business objectives, and ensure cloud environments are secure, compliant, and scalable, without requiring hands on coding. Technology stack: Google Cloud Platform (GCP), AWS, Azure, GCP Cloud Adoption Framework, IAM, VPC Service Controls, Organization Policies, Databricks, PostgreSQL, dbt, Snowflake, Apigee, CI/CD pipelines, and FinOps practices. Responsibilities Define scalable and resilient cloud architectures aligned with enterprise standards Collaborate with engineering teams to design cloud solutions using structured adoption frameworks Translate business requirements into clear architecture blueprints Develop migration strategies, roadmaps, and cost aware planning models Communicate architectural decisions, risks, and progress to stakeholders in a clear manner Ensure alignment with enterprise governance, compliance, and security standards Identify and mitigate technical and operational risks early in the lifecycle Act as a bridge between business stakeholders, compliance teams, and engineering groups Guide adoption of modern cloud architecture patterns and best practices Support improvements in cloud maturity, scalability, and operational efficiency Requirements Experience working as a Lead or Solution Architect in complex enterprise environments Strong experience designing cloud architectures and infrastructure solutions Knowledge of Google Cloud Platform including landing zones, resource hierarchy, and organizational structure Understanding of cloud security principles including identity management and governance frameworks Experience defining high level architecture blueprints aligned with business strategies Knowledge of data governance practices including data lifecycle management and access control Understanding of regulatory and compliance requirements in financial environments Experience working with multi cloud environments including AWS and Azure Familiarity with tools such as Databricks, PostgreSQL, dbt, Snowflake, and Apigee Strong communication and stakeholder management skillsExperience working within Agile and DevOps environments Nice to have Experience with FinOps practices and cloud cost optimization strategies Exposure to hybrid cloud implementations Experience supporting enterprise scale cloud migration programs Familiarity with observability tools and cloud monitoring solutions
This role is hybrid, UK-based. About the role This is an exciting opportunity to lead the security, resilience and operational integrity of EA Technology's cloud platforms and live software applications. You will play a critical role in ensuring our SaaS platforms operate securely, reliably and at scale, supporting high availability in safety-critical environments. Responsibilities Own end-to-end platform resilience, ensuring SaaS systems and cloud services meet availability, performance and recovery objectives Define and enforce secure by design principles, including cybersecurity standards, cloud architecture guardrails and operational controls Lead DevOps and Site Reliability Engineering (SRE) maturity, embedding monitoring, observability, automated testing and structured incident response Drive adoption of automation and AI enabled tooling to improve anomaly detection, incident management, vulnerability management and operational efficiency Provide technical oversight of production operations, ensuring alignment between platform standards and operational support teams Lead cybersecurity governance, including audits, penetration testing, vulnerability management and risk frameworks Support integration of acquired platforms, ensuring alignment with EA Technology's security and reliability standards Establish and maintain platform standards that balance delivery speed with system integrity and resilience What we'll need from you Significant experience leading platform engineering, cybersecurity or reliability teams in complex SaaS environments Deep expertise in cloud native architecture (Azure, AWS or equivalent) Strong understanding of DevOps, SRE principles and SaaS production operations Experience implementing secure by design frameworks and managing cybersecurity governance Experience owning uptime, reliability and systemic operational performance Strong leadership skills, with the ability to work across Engineering, Product and Operations Excellent communication skills, with the ability to articulate technical risk and trade offs clearly Ability to lead calmly and effectively during high pressure incidents Desirable Experience in utilities, energy or critical infrastructure sectors Experience integrating acquired digital platforms Familiarity with hybrid IT/OT security models Experience with FinOps and cloud cost optimisation practices Experience applying AI assisted tooling in operational or security environments What we can offer you Competitive salary + company performance bonus Career development opportunities: We offer genuine pathways for growth within our company Work life balance: With flexible working options, we support our employees in balancing their professional and personal lives Holidays: 25 days of annual leave, plus bank holidays Pension contributions of 8% from the employer (or cash equivalent) Comprehensive benefits, including Group Life Insurance, Income Protection, and Critical Illness cover (or cash equivalents) Private Medical Insurance (single cover or cash equivalent) A truly collaborative and supportive work environment where amazing colleagues inspire each other every day!
27/07/2026
Full time
This role is hybrid, UK-based. About the role This is an exciting opportunity to lead the security, resilience and operational integrity of EA Technology's cloud platforms and live software applications. You will play a critical role in ensuring our SaaS platforms operate securely, reliably and at scale, supporting high availability in safety-critical environments. Responsibilities Own end-to-end platform resilience, ensuring SaaS systems and cloud services meet availability, performance and recovery objectives Define and enforce secure by design principles, including cybersecurity standards, cloud architecture guardrails and operational controls Lead DevOps and Site Reliability Engineering (SRE) maturity, embedding monitoring, observability, automated testing and structured incident response Drive adoption of automation and AI enabled tooling to improve anomaly detection, incident management, vulnerability management and operational efficiency Provide technical oversight of production operations, ensuring alignment between platform standards and operational support teams Lead cybersecurity governance, including audits, penetration testing, vulnerability management and risk frameworks Support integration of acquired platforms, ensuring alignment with EA Technology's security and reliability standards Establish and maintain platform standards that balance delivery speed with system integrity and resilience What we'll need from you Significant experience leading platform engineering, cybersecurity or reliability teams in complex SaaS environments Deep expertise in cloud native architecture (Azure, AWS or equivalent) Strong understanding of DevOps, SRE principles and SaaS production operations Experience implementing secure by design frameworks and managing cybersecurity governance Experience owning uptime, reliability and systemic operational performance Strong leadership skills, with the ability to work across Engineering, Product and Operations Excellent communication skills, with the ability to articulate technical risk and trade offs clearly Ability to lead calmly and effectively during high pressure incidents Desirable Experience in utilities, energy or critical infrastructure sectors Experience integrating acquired digital platforms Familiarity with hybrid IT/OT security models Experience with FinOps and cloud cost optimisation practices Experience applying AI assisted tooling in operational or security environments What we can offer you Competitive salary + company performance bonus Career development opportunities: We offer genuine pathways for growth within our company Work life balance: With flexible working options, we support our employees in balancing their professional and personal lives Holidays: 25 days of annual leave, plus bank holidays Pension contributions of 8% from the employer (or cash equivalent) Comprehensive benefits, including Group Life Insurance, Income Protection, and Critical Illness cover (or cash equivalents) Private Medical Insurance (single cover or cash equivalent) A truly collaborative and supportive work environment where amazing colleagues inspire each other every day!
Cloud Engineering ManagerLondon Council Permanent Salary: £79,005 to £88,149 Location: Westminster, London Hybrid and Flexible Working The Opportunity Salt is exclusively partnering with a London based Council to recruit a Cloud Engineering Manager to lead a highly skilled Cloud Engineering team through one of the council's most exciting technology transformation programmes. This is a fantastic opportunity to build and shape modern Azure and Google Cloud environments from the ground up rather than simply maintaining Legacy infrastructure. You'll lead a team of up to six Cloud Engineers while remaining heavily involved in technical strategy, solution design and the delivery of large scale cloud transformation projects. Approximately 20% of the role focuses on leadership and people management, with the remaining 80% dedicated to technical leadership, cloud architecture, engineering strategy and delivering modern cloud platforms. Responsibilities Lead, coach and develop a team of up to six Cloud Engineers. Manage Agile delivery, sprint planning, stand ups and engineering ceremonies. Design and implement secure Azure and Google Cloud environments. Deliver Azure Landing Zones and modern cloud platforms. Drive Zero Trust cloud architecture and security best practice. Implement Infrastructure as Code using Terraform and Bicep. Apply Microsoft Cloud Adoption Framework (CAF) principles. Lead cloud migration and modernisation initiatives. Drive DevOps and Site Reliability Engineering (SRE) practices. Ensure cloud governance, resilience and operational excellence. Work closely with Architecture, Security, Development, Data and Infrastructure teams. Review new project requests and provide technical direction, solution design and delivery estimates. Manage Azure cost optimisation and FinOps activities. Build delivery roadmaps and continually improve engineering processes. Manage supplier relationships and technical delivery partners. Essential Experience Experience leading Cloud Engineering or Infrastructure Engineering teams. Strong Azure engineering and architecture experience. Experience working with Google Cloud Platform (GCP). Azure Landing Zones. Azure networking, identity and cloud security. Terraform and Infrastructure as Code. Bicep. Azure DevOps. DevOps and CI/CD. Site Reliability Engineering (SRE). Cloud governance. Cloud migration programmes. Azure cost optimisation and FinOps. Technical solution design. Agile, Scrum and Kanban. Stakeholder management. Supplier management. Excellent communication and leadership skills. Desirable Local Government experience. Microsoft Cloud Adoption Framework (CAF). Zero Trust architecture. Multi cloud environments. Enterprise scale transformation programmes. Benefits Salary up to £88,149. Hybrid and flexible working. 36 hour working week. 31 days annual leave plus your birthday off. Excellent Local Government Pension Scheme. Private Medical Insurance. Cycle to Work Scheme. Employee discounts. Learning and development opportunities. This is an outstanding opportunity for an experienced Cloud Engineering leader who enjoys remaining technically involved whilst leading the delivery of enterprise scale cloud transformation across one of London's leading local authorities. Rates depend on experience and client requirements
27/07/2026
Full time
Cloud Engineering ManagerLondon Council Permanent Salary: £79,005 to £88,149 Location: Westminster, London Hybrid and Flexible Working The Opportunity Salt is exclusively partnering with a London based Council to recruit a Cloud Engineering Manager to lead a highly skilled Cloud Engineering team through one of the council's most exciting technology transformation programmes. This is a fantastic opportunity to build and shape modern Azure and Google Cloud environments from the ground up rather than simply maintaining Legacy infrastructure. You'll lead a team of up to six Cloud Engineers while remaining heavily involved in technical strategy, solution design and the delivery of large scale cloud transformation projects. Approximately 20% of the role focuses on leadership and people management, with the remaining 80% dedicated to technical leadership, cloud architecture, engineering strategy and delivering modern cloud platforms. Responsibilities Lead, coach and develop a team of up to six Cloud Engineers. Manage Agile delivery, sprint planning, stand ups and engineering ceremonies. Design and implement secure Azure and Google Cloud environments. Deliver Azure Landing Zones and modern cloud platforms. Drive Zero Trust cloud architecture and security best practice. Implement Infrastructure as Code using Terraform and Bicep. Apply Microsoft Cloud Adoption Framework (CAF) principles. Lead cloud migration and modernisation initiatives. Drive DevOps and Site Reliability Engineering (SRE) practices. Ensure cloud governance, resilience and operational excellence. Work closely with Architecture, Security, Development, Data and Infrastructure teams. Review new project requests and provide technical direction, solution design and delivery estimates. Manage Azure cost optimisation and FinOps activities. Build delivery roadmaps and continually improve engineering processes. Manage supplier relationships and technical delivery partners. Essential Experience Experience leading Cloud Engineering or Infrastructure Engineering teams. Strong Azure engineering and architecture experience. Experience working with Google Cloud Platform (GCP). Azure Landing Zones. Azure networking, identity and cloud security. Terraform and Infrastructure as Code. Bicep. Azure DevOps. DevOps and CI/CD. Site Reliability Engineering (SRE). Cloud governance. Cloud migration programmes. Azure cost optimisation and FinOps. Technical solution design. Agile, Scrum and Kanban. Stakeholder management. Supplier management. Excellent communication and leadership skills. Desirable Local Government experience. Microsoft Cloud Adoption Framework (CAF). Zero Trust architecture. Multi cloud environments. Enterprise scale transformation programmes. Benefits Salary up to £88,149. Hybrid and flexible working. 36 hour working week. 31 days annual leave plus your birthday off. Excellent Local Government Pension Scheme. Private Medical Insurance. Cycle to Work Scheme. Employee discounts. Learning and development opportunities. This is an outstanding opportunity for an experienced Cloud Engineering leader who enjoys remaining technically involved whilst leading the delivery of enterprise scale cloud transformation across one of London's leading local authorities. Rates depend on experience and client requirements
The TP ICAP Group is a world leading provider of market infrastructure.Our purpose is to provide clients with access to global financial and commodities markets, improving price discovery, liquidity, and distribution of data, through responsible and innovative solutions.Through our people and technology, we connect clients to superior liquidity and data solutions.The Group is home to a stable of premium brands. Collectively, TP ICAP is the largest interdealer broker in the world by revenue, the number one Energy & Commodities broker in the world, the world's leading provider of OTC data, and an award winning all-to-all trading platform.The Group operates from more than 48 offices in 27 countries. We are 5,200 people strong. We work as one to achieve our vision of being the world's most trusted, innovative, liquidity and data solutions specialist.Role Overview:The Procurement function is responsible for the provision of leading edge procurement services aligned to business strategy and requirements, aiming to enhance value and productivity across the supplier value-chain through effective, efficient and agile processes delivered by a strategic, commercial and risk focussed function.The function plays a crucial role in ensuring that the organisation achieves value (i.e. performance, financial, enhanced risk, etc.) from its global supplier landscape. As such Procurement is accountable for the implementation and management of strategic procurement driving value through effective supplier relationship management, demand & consumption management, MI & BI analytics, third party risk management and the like.Please note, this role is a 12-month Fixed Term Contract. The successful candidate will provide specialist Technology Procurement expertise to support strategic sourcing initiatives, commercial optimisation activities, supplier governance and procurement transformation programmes across the Group during this period.Role Responsibilities:Support the Head of Procurement - Technology with development and implementation of 3 year rolling Technology Category Sourcing Strategy, including supplier segmentation, supplier consolidation strategy, PSL category definition and strategic procurement levers implementationDefine and agree category plans that deliver ongoing commercial value in terms of cost, performance and risk in accordance with the overall procurement strategy aligned to the business objectives and strategiesBuild strong relationships with technology and business stakeholders, business management and finance as well ask third party suppliersProvide market insights to the broader procurement team including key stakeholders, senior management and the procurement sourcing officeDeliver sourcing process including pipeline management, competitive tenders and benchmarking to identify best supplier option is identified and commercial, contractual, performance and service delivery stipulations are embedded and aligned to business requirements.Ensure and complete all relevant Third Party Risk Management activities on a timely basis through liaison with the Procurement Efficiency OfficeEnsure that all technology software contracts have been abstracted and added to the contract repository and contract meta data is kept up-to-date throughout the contract lifecycle.Actively participate in the Procurement strategic initiatives including transformation, continuous improvement and innovation programmesManage and lead complex and high value contract negotiations with suppliersParticipate and manage appropriate Governance forums and develop supporting metrics & reporting in conjunction with the Procurement Efficiency Office.Collaborate and support the commercial contract management team with the management of the category Panels and Preferred Supplier Lists, performance and commercial issues during the post-deal contract period.Provide support to the Procurement Efficiency Office with supplier onboarding queries and process issues, as well as offboarding of suppliers upon contract termination.Experience / CompetencesEssentialProven experience operating within a strategic procurement, category management or sourcing function.Demonstrable experience managing Technology spend categories, including one or more of the following: Software, SaaS, Infrastructure, Cloud Services, Telecommunications, Managed Services or IT Professional Services.Experience leading end-to-end sourcing activities, including supplier selection, competitive tendering, negotiations and contract award.Proven commercial negotiation and contract management skills, with the ability to deliver measurable value through cost optimisation, risk reduction and service improvements.Experience developing and executing category plans or procurement strategies aligned to business objectives.Demonstrated stakeholder management and relationship-building skills, with the ability to influence stakeholders across Technology, Finance and Business functions.Experience managing supplier relationships and evaluating supplier performance, capability and commercial value.Solid analytical skills with the ability to interpret spend data, market intelligence and supplier insights to support decision making.Excellent verbal and written communication skills, including the ability to present recommendations and commercial outcomes to senior stakeholders.Proven ability to manage multiple priorities and deliver outcomes within a fast-paced and evolving environment.DesiredExperience managing Software Licensing, SaaS, Cloud Infrastructure or IT Outsourcing categories in a complex enterprise environment.Experience within Financial Services, Capital Markets, Professional Services or another regulated industry.Knowledge of cloud commercial models, FinOps principles and cloud cost optimisation initiatives.Experience with Third-Party Risk Management, supplier governance and regulatory requirements relating to external suppliers.Familiarity with technology-related contracts, including software licensing, SaaS agreements, managed services and professional services contracts.Experience supporting procurement transformation, process improvement or operating model enhancement initiatives.Professional procurement qualification (MCIPS or equivalent).Experience working within a global procurement function supporting international stakeholders and suppliers.Experience using procurement, contract lifecycle management and supplier management platforms.Band / LevelManager / 7 The Perfect Fit?Concerned that you may not meet the criteria precisely? At TP ICAP, we wholeheartedly believe in fostering inclusivity and cultivating a work environment where everyone can flourish, regardless of your personal or professional background. If you are enthusiastic about this role but find that your experience doesn't align perfectly with every aspect of the job description, we strongly encourage you to apply. You may be the ideal candidate for this position or another opportunity within our organisation. Our dedicated Talent Acquisition team is here to assist you in recognising how your unique skills and abilities can be a valuable contribution. Don't hesitate to take the leap and explore the possibilities. Your potential is what truly matters to us.Company StatementWe know that the best innovation happens when diverse people with different perspectives and skills work together in an inclusive atmosphere. That's why we're building a culture where everyone plays a part in making people feel welcome, ready and willing to contribute. TP ICAP Accord - our Employee Network - is a central to this. As well as representing specific groups, TP ICAP Accord helps increase awareness, collaboration, shares best practice, and holds our firm to account for driving continuous cultural improvement.LocationUK - 135 Bishopsgate - London
27/07/2026
Full time
The TP ICAP Group is a world leading provider of market infrastructure.Our purpose is to provide clients with access to global financial and commodities markets, improving price discovery, liquidity, and distribution of data, through responsible and innovative solutions.Through our people and technology, we connect clients to superior liquidity and data solutions.The Group is home to a stable of premium brands. Collectively, TP ICAP is the largest interdealer broker in the world by revenue, the number one Energy & Commodities broker in the world, the world's leading provider of OTC data, and an award winning all-to-all trading platform.The Group operates from more than 48 offices in 27 countries. We are 5,200 people strong. We work as one to achieve our vision of being the world's most trusted, innovative, liquidity and data solutions specialist.Role Overview:The Procurement function is responsible for the provision of leading edge procurement services aligned to business strategy and requirements, aiming to enhance value and productivity across the supplier value-chain through effective, efficient and agile processes delivered by a strategic, commercial and risk focussed function.The function plays a crucial role in ensuring that the organisation achieves value (i.e. performance, financial, enhanced risk, etc.) from its global supplier landscape. As such Procurement is accountable for the implementation and management of strategic procurement driving value through effective supplier relationship management, demand & consumption management, MI & BI analytics, third party risk management and the like.Please note, this role is a 12-month Fixed Term Contract. The successful candidate will provide specialist Technology Procurement expertise to support strategic sourcing initiatives, commercial optimisation activities, supplier governance and procurement transformation programmes across the Group during this period.Role Responsibilities:Support the Head of Procurement - Technology with development and implementation of 3 year rolling Technology Category Sourcing Strategy, including supplier segmentation, supplier consolidation strategy, PSL category definition and strategic procurement levers implementationDefine and agree category plans that deliver ongoing commercial value in terms of cost, performance and risk in accordance with the overall procurement strategy aligned to the business objectives and strategiesBuild strong relationships with technology and business stakeholders, business management and finance as well ask third party suppliersProvide market insights to the broader procurement team including key stakeholders, senior management and the procurement sourcing officeDeliver sourcing process including pipeline management, competitive tenders and benchmarking to identify best supplier option is identified and commercial, contractual, performance and service delivery stipulations are embedded and aligned to business requirements.Ensure and complete all relevant Third Party Risk Management activities on a timely basis through liaison with the Procurement Efficiency OfficeEnsure that all technology software contracts have been abstracted and added to the contract repository and contract meta data is kept up-to-date throughout the contract lifecycle.Actively participate in the Procurement strategic initiatives including transformation, continuous improvement and innovation programmesManage and lead complex and high value contract negotiations with suppliersParticipate and manage appropriate Governance forums and develop supporting metrics & reporting in conjunction with the Procurement Efficiency Office.Collaborate and support the commercial contract management team with the management of the category Panels and Preferred Supplier Lists, performance and commercial issues during the post-deal contract period.Provide support to the Procurement Efficiency Office with supplier onboarding queries and process issues, as well as offboarding of suppliers upon contract termination.Experience / CompetencesEssentialProven experience operating within a strategic procurement, category management or sourcing function.Demonstrable experience managing Technology spend categories, including one or more of the following: Software, SaaS, Infrastructure, Cloud Services, Telecommunications, Managed Services or IT Professional Services.Experience leading end-to-end sourcing activities, including supplier selection, competitive tendering, negotiations and contract award.Proven commercial negotiation and contract management skills, with the ability to deliver measurable value through cost optimisation, risk reduction and service improvements.Experience developing and executing category plans or procurement strategies aligned to business objectives.Demonstrated stakeholder management and relationship-building skills, with the ability to influence stakeholders across Technology, Finance and Business functions.Experience managing supplier relationships and evaluating supplier performance, capability and commercial value.Solid analytical skills with the ability to interpret spend data, market intelligence and supplier insights to support decision making.Excellent verbal and written communication skills, including the ability to present recommendations and commercial outcomes to senior stakeholders.Proven ability to manage multiple priorities and deliver outcomes within a fast-paced and evolving environment.DesiredExperience managing Software Licensing, SaaS, Cloud Infrastructure or IT Outsourcing categories in a complex enterprise environment.Experience within Financial Services, Capital Markets, Professional Services or another regulated industry.Knowledge of cloud commercial models, FinOps principles and cloud cost optimisation initiatives.Experience with Third-Party Risk Management, supplier governance and regulatory requirements relating to external suppliers.Familiarity with technology-related contracts, including software licensing, SaaS agreements, managed services and professional services contracts.Experience supporting procurement transformation, process improvement or operating model enhancement initiatives.Professional procurement qualification (MCIPS or equivalent).Experience working within a global procurement function supporting international stakeholders and suppliers.Experience using procurement, contract lifecycle management and supplier management platforms.Band / LevelManager / 7 The Perfect Fit?Concerned that you may not meet the criteria precisely? At TP ICAP, we wholeheartedly believe in fostering inclusivity and cultivating a work environment where everyone can flourish, regardless of your personal or professional background. If you are enthusiastic about this role but find that your experience doesn't align perfectly with every aspect of the job description, we strongly encourage you to apply. You may be the ideal candidate for this position or another opportunity within our organisation. Our dedicated Talent Acquisition team is here to assist you in recognising how your unique skills and abilities can be a valuable contribution. Don't hesitate to take the leap and explore the possibilities. Your potential is what truly matters to us.Company StatementWe know that the best innovation happens when diverse people with different perspectives and skills work together in an inclusive atmosphere. That's why we're building a culture where everyone plays a part in making people feel welcome, ready and willing to contribute. TP ICAP Accord - our Employee Network - is a central to this. As well as representing specific groups, TP ICAP Accord helps increase awareness, collaboration, shares best practice, and holds our firm to account for driving continuous cultural improvement.LocationUK - 135 Bishopsgate - London
Hitachi Digital Services seeks a senior FinOps leader to own the on prem FinOps strategy, cost governance, and data architecture, driving visibility and cost accountability across executives, finance, engineering, and operations. You will define data models, tagging standards, and reporting ecosystems to enable proactive cost optimisation and transparent budgeting for the organization's on premises estate.
27/07/2026
Full time
Hitachi Digital Services seeks a senior FinOps leader to own the on prem FinOps strategy, cost governance, and data architecture, driving visibility and cost accountability across executives, finance, engineering, and operations. You will define data models, tagging standards, and reporting ecosystems to enable proactive cost optimisation and transparent budgeting for the organization's on premises estate.
About the Role Networks Manager - EMEA (Network Service Owner) About the Role We're looking for an experienced Networks Manager - EMEA to to lead the strategy, performance and evolution of the network services underpinning our EMEA wide business. Reporting to the Infrastructure Director, this is a critical leadership role within the Infrastructure, Security and Services organisation, responsible for ensuring the delivery of secure, resilient and high-performing network services that meet the needs of a complex international business. You'll combine operational excellence with strategic thinking, driving service quality today while contributing to shaping the future network roadmap alongside our team of architects. As the operational owner of network services, you'll coordinate and manage third-party partners, ensuring robust service management across the scope of our network services. You'll play a role in influencing technology decisions, supporting business transformation initiatives and ensuring our network capabilities continue to evolve in line with business needs. Key Responsibilities Own and lead the end-to-end network service, ensuring availability, capacity and performance targets are consistently achieved Support definition of the network technology roadmap, balancing operational stability, innovation and future business requirements Delivery of network services through 3rd party vendors, leveraging data and dashboards to understand service and vendor performance and to hold suppliers to account as needed Drive operational excellence through effective service management, including incident, problem, change, risk and supplier management Partner closely with IT Business Partners and stakeholders to ensure service delivery meets business expectations Lead the successful transition of new network services and technologies into operational support through alignment with delivery teams Establish and maintain standards, policies and procedures aligned to industry best practice Manage service risks and collaborate with governance teams to ensure appropriate visibility and mitigation plans are in place Develop strong FinOps disciplines, ensuring effective cost management and budget control across the network estate Provide technical leadership and subject matter expertise, enabling informed decision-making across Infrastructure and Technology teams Support major projects with resource planning and support from within the supplier landscape Coordinate and manage third-party suppliers to ensure high-quality service delivery and value for money, working with our supplier management function and suppliers as needed to implement service improvement plans Drive continuous improvement initiatives, leveraging operational data and service insights to enhance performance About You You're an experienced infrastructure or network leader who thrives in complex enterprise environments and combines strong technical expertise with excellent stakeholder management skills. We're looking for: Significant experience leading enterprise network services within a large, complex organisation A highly operational focus, with prior experience defining and monitoring day to day practices which proactively protect the service Experience of major incident management within a business critical service - managing stakeholder expectations whilst driving resolution through technical 3rd parties Strong knowledge of network technologies, operations and service management best practices Experience managing network vendors, suppliers and outsourced service providers within a 100% outsourced model Demonstrated success in transforming outsourced network functions to deliver operational excellence across critical infrastructure services Strong understanding of ITIL disciplines, including Incident, Problem, Change and Service Management Proven ability to influence senior stakeholders and communicate complex technical topics to non-technical audiences A clear viewpoint on effective service reporting and governance with a proven ability to leverage that to deliver meaningful results Strong commercial awareness, including budget management, cost control and FinOps principles Strong analytical and problem-solving skills, with the ability to make informed decisions in fast-paced environments We are DS Smith, together with International Paper, we are a global leader in sustainable packaging solutions and other fibre-based products.We believe a better, more sustainable tomorrow is possible with the right people, who challenge and support one another to enact positive change. We employ more than 60,000 colleagues in North America and Europe, Middle East and Africa (EMEA), who are experts in innovation,manufacturing, design, sales, sustainability, supply chain, and much more. Together with our customers, we make the world safer and more productive, one sustainable packaging solution at a time. Become part of a world-leading organisation and do your best work with us! As the journey continues of bringing together the strengths of both organisations, during your candidate experience you may engage with our colleagues from International Paper! You could visit an International Paper or DS Smith site or office.
27/07/2026
Full time
About the Role Networks Manager - EMEA (Network Service Owner) About the Role We're looking for an experienced Networks Manager - EMEA to to lead the strategy, performance and evolution of the network services underpinning our EMEA wide business. Reporting to the Infrastructure Director, this is a critical leadership role within the Infrastructure, Security and Services organisation, responsible for ensuring the delivery of secure, resilient and high-performing network services that meet the needs of a complex international business. You'll combine operational excellence with strategic thinking, driving service quality today while contributing to shaping the future network roadmap alongside our team of architects. As the operational owner of network services, you'll coordinate and manage third-party partners, ensuring robust service management across the scope of our network services. You'll play a role in influencing technology decisions, supporting business transformation initiatives and ensuring our network capabilities continue to evolve in line with business needs. Key Responsibilities Own and lead the end-to-end network service, ensuring availability, capacity and performance targets are consistently achieved Support definition of the network technology roadmap, balancing operational stability, innovation and future business requirements Delivery of network services through 3rd party vendors, leveraging data and dashboards to understand service and vendor performance and to hold suppliers to account as needed Drive operational excellence through effective service management, including incident, problem, change, risk and supplier management Partner closely with IT Business Partners and stakeholders to ensure service delivery meets business expectations Lead the successful transition of new network services and technologies into operational support through alignment with delivery teams Establish and maintain standards, policies and procedures aligned to industry best practice Manage service risks and collaborate with governance teams to ensure appropriate visibility and mitigation plans are in place Develop strong FinOps disciplines, ensuring effective cost management and budget control across the network estate Provide technical leadership and subject matter expertise, enabling informed decision-making across Infrastructure and Technology teams Support major projects with resource planning and support from within the supplier landscape Coordinate and manage third-party suppliers to ensure high-quality service delivery and value for money, working with our supplier management function and suppliers as needed to implement service improvement plans Drive continuous improvement initiatives, leveraging operational data and service insights to enhance performance About You You're an experienced infrastructure or network leader who thrives in complex enterprise environments and combines strong technical expertise with excellent stakeholder management skills. We're looking for: Significant experience leading enterprise network services within a large, complex organisation A highly operational focus, with prior experience defining and monitoring day to day practices which proactively protect the service Experience of major incident management within a business critical service - managing stakeholder expectations whilst driving resolution through technical 3rd parties Strong knowledge of network technologies, operations and service management best practices Experience managing network vendors, suppliers and outsourced service providers within a 100% outsourced model Demonstrated success in transforming outsourced network functions to deliver operational excellence across critical infrastructure services Strong understanding of ITIL disciplines, including Incident, Problem, Change and Service Management Proven ability to influence senior stakeholders and communicate complex technical topics to non-technical audiences A clear viewpoint on effective service reporting and governance with a proven ability to leverage that to deliver meaningful results Strong commercial awareness, including budget management, cost control and FinOps principles Strong analytical and problem-solving skills, with the ability to make informed decisions in fast-paced environments We are DS Smith, together with International Paper, we are a global leader in sustainable packaging solutions and other fibre-based products.We believe a better, more sustainable tomorrow is possible with the right people, who challenge and support one another to enact positive change. We employ more than 60,000 colleagues in North America and Europe, Middle East and Africa (EMEA), who are experts in innovation,manufacturing, design, sales, sustainability, supply chain, and much more. Together with our customers, we make the world safer and more productive, one sustainable packaging solution at a time. Become part of a world-leading organisation and do your best work with us! As the journey continues of bringing together the strengths of both organisations, during your candidate experience you may engage with our colleagues from International Paper! You could visit an International Paper or DS Smith site or office.
We're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people-centric and here to power good.Every day, we future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our companyand customers from what's now to what's next.We make it happen through the power of acceleration. Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us. Job Description Mandatory Skills: Certified from FinOps.org, FinOps for DC knowledge, DCIM, CMDB / asset DB, hypervisor metrics, power monitoring, contracts & licenses management Role Description Skills: ROLE PURPOSE Own the end-to-end FinOps strategy and architecture for the organisation's on-premises estate, following the FinOps Foundation (FinOps.org) framework across the Inform, Optimise, and Operate lifecycle phases. Define the cost visibility strategy, establish the data model and governance framework, lead maturity assessments, and architect the reporting and dashboard ecosystem. Serve as the primary FinOps authority, engaging all FinOps personas - from executives and finance through engineering and operations - to build a culture of cost accountability and data-driven infrastructure economics. This role bridges the gap between financial management and engineering, translating on-premises infrastructure complexity into clear cost models, unit economics, and actionable savings roadmaps. FINOPS SCOPE OF WORK This role directly contributes to the following FinOps programme workstreams, aligned to FinOps.org's Inform, Optimise, and Operate lifecycle: Cost Visibility - Gain full visibility of on-premises infrastructure, licensing, and operational costs by establishing comprehensive data collection and integration from identified sources Data Source Identification & Setup - Identify, agree upon, and integrate all relevant cost data sources including CMDB, asset management, licensing portals, procurement systems, and infrastructure monitoring tools FinOps Maturity Baselining - Assess current FinOps maturity against the FinOps.org maturity model (Crawl -> Walk -> Run), identifying gaps and opportunities across tagging, savings, and cost governance Tagging & Allocation Gaps - Identify and remediate gaps in cost allocation, tagging, and labelling across on-premises assets to enable accurate chargeback, showback, and cost attribution Savings Identification - Discover potential savings opportunities through licence optimisation, hardware rightsizing, capacity rebalancing, decommissioning, and vendor renegotiation Cost Governance Gaps - Assess and strengthen cost governance processes including approval workflows, budget controls, anomaly detection, and policy enforcement Data Normalisation - Normalise cost data from disparate sources into a unified cost model, enabling consistent reporting across business units, environments, and service categories Usage, Pattern & Licence Analysis - Analyse infrastructure utilisation patterns, workload distribution, peak/off-peak trends, and software licence usage to identify waste and optimisation opportunities Cost Modelling & Economics - Define unit economics, cost-per-service models, TCO calculations, and amortisation schedules for on-premises infrastructure to enable informed decision-making Implementation Roadmap - Develop a prioritised roadmap of quick-wins (0-3 months), mid-term initiatives (3-6 months), and long-term transformations (6-12 months) based on value and feasibility Reporting & Dashboards - Design and recommend live/near-real-time cost dashboards with drill-down capabilities, trend analysis, anomaly alerts, and executive reporting views FINOPS PERSONA COLLABORATION As defined by FinOps.org, this role works with all key FinOps personas: Executives (CTO, CFO, CIO) - Present cost visibility dashboards, business cases for optimisation, and FinOps maturity roadmaps to drive strategic investment decisions Finance / Procurement - Collaborate on cost allocation models, chargeback/showback mechanisms, budgeting, forecasting, and vendor contract analysis for on-premises licensing Engineering / DevOps Teams - Partner on tagging strategies, resource utilisation analysis, rightsizing recommendations, and embedding cost-awareness into engineering culture Product Owners / Business Units - Provide unit economics and cost-per-service metrics to inform product decisions, capacity planning, and build-vs-buy analysis IT Operations / Infrastructure - Work on data source integration, asset management, capacity utilisation, and hardware lifecycle cost modelling Architecture & Governance - Align FinOps governance policies, cost guardrails, and architectural decisions with total cost of ownership (TCO) analysis Security & Compliance - Ensure cost data handling meets compliance requirements and that FinOps tooling adheres to organisational security policies KEY RESPONSIBILITIES Define and own the on-premises FinOps strategy, architecture, and maturity roadmap aligned to the FinOps.org Inform -> Optimise -> Operate lifecycle Lead the FinOps maturity assessment (Crawl/Walk/Run) across tagging, savings identification, and cost governance - produce a baseline scorecard with prioritised improvement actions Architect the cost data model: identify and integrate all data sources (CMDB, asset management, procurement, licensing, monitoring), define normalisation rules, and establish the single source of truth for on-premises cost data Design cost allocation frameworks: define chargeback/showback models, tagging taxonomies, and cost attribution rules that enable accurate cost-per-business-unit, cost-per-service, and cost-per-environment reporting Develop unit economics and cost modelling: build TCO models, amortisation schedules, build-vs-buy comparisons, and cost-per-transaction metrics for on-premises infrastructure Lead savings discovery workshops across all FinOps personas: licence optimisation (true-ups, shelfware, harvesting), hardware rightsizing, capacity rebalancing, and decommissioning opportunities Architect the FinOps reporting and dashboard ecosystem: define requirements for live/near-real-time dashboards with executive views, drill-down capabilities, trend analysis, anomaly detection, and budget variance alerts Establish FinOps governance framework: cost approval policies, budget guardrails, anomaly response procedures, and regular cost review cadences with all persona groups Create the prioritised FinOps implementation roadmap: quick-wins (0-3 months), mid-term initiatives (3-6 months), and long-term transformation (6-12 months) based on value, feasibility, and stakeholder alignment Mentor and lead the FinOps engineering team (1 onshore + 1 offshore), establish engineering standards, and drive a culture of cost transparency and accountability across the organisation Present FinOps insights, savings opportunities, and governance reports to executive stakeholders (CTO, CFO, CIO) and lead regular FinOps review forums Stay current with FinOps.org framework evolution, contribute to the FinOps community of practice, and continuously refine methodologies for on-premises environments TECHNICAL SKILLS & EXPERTISE Deep expertise in on-premises infrastructure economics: server, storage, network, licensing (Microsoft EA, Oracle, VMware, Red Hat), data centre costs (power, cooling, space), and depreciation/amortisation models Strong understanding of the FinOps.org framework: Inform, Optimise, Operate lifecycle; FinOps Domains (Cost Allocation, Anomaly Management, Rate Optimisation, Usage Optimisation, etc.); FinOps maturity model (Crawl/Walk/Run) Experience with cost data integration from: CMDB (ServiceNow), asset management tools, procurement platforms, licence management tools (Snow, Flexera), and infrastructure monitoring (Datadog, Dynatrace, SolarWinds) Data modelling and analytics: SQL, Python/Pandas for cost data analysis; experience building normalisation pipelines and cost allocation engines Dashboard and reporting tools: Power BI, Tableau, Grafana, or custom BI solutions for real-time cost visibility and executive dashboards Understanding of ITIL asset and configuration management, hardware lifecycle management, and capacity planning methodologies Familiarity with FinOps tooling: Apptio, CloudHealth/VMware Aria, Flexera One, ServiceNow ITOM/ITFM, or custom-built cost management platforms Knowledge of hybrid/multi-cloud cost management to bridge on-premises FinOps with cloud FinOps practices where applicable SOFT SKILLS & COMPETENCIES Exceptional stakeholder management - able to influence executives, finance, engineering, and operations teams with equal credibility Strategic thinker who translates complex cost data into clear business narratives and actionable recommendations Strong leadership and mentoring ability - coaches and develops FinOps engineers and builds a cost-aware engineering culture Excellent presentation and communication skills - comfortable presenting to C-level audiences and leading governance forums Commercial acumen - understands P&L, depreciation, amortisation, OpEx/CapEx distinctions, and vendor contract structures . click apply for full job details
27/07/2026
Full time
We're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people-centric and here to power good.Every day, we future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our companyand customers from what's now to what's next.We make it happen through the power of acceleration. Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us. Job Description Mandatory Skills: Certified from FinOps.org, FinOps for DC knowledge, DCIM, CMDB / asset DB, hypervisor metrics, power monitoring, contracts & licenses management Role Description Skills: ROLE PURPOSE Own the end-to-end FinOps strategy and architecture for the organisation's on-premises estate, following the FinOps Foundation (FinOps.org) framework across the Inform, Optimise, and Operate lifecycle phases. Define the cost visibility strategy, establish the data model and governance framework, lead maturity assessments, and architect the reporting and dashboard ecosystem. Serve as the primary FinOps authority, engaging all FinOps personas - from executives and finance through engineering and operations - to build a culture of cost accountability and data-driven infrastructure economics. This role bridges the gap between financial management and engineering, translating on-premises infrastructure complexity into clear cost models, unit economics, and actionable savings roadmaps. FINOPS SCOPE OF WORK This role directly contributes to the following FinOps programme workstreams, aligned to FinOps.org's Inform, Optimise, and Operate lifecycle: Cost Visibility - Gain full visibility of on-premises infrastructure, licensing, and operational costs by establishing comprehensive data collection and integration from identified sources Data Source Identification & Setup - Identify, agree upon, and integrate all relevant cost data sources including CMDB, asset management, licensing portals, procurement systems, and infrastructure monitoring tools FinOps Maturity Baselining - Assess current FinOps maturity against the FinOps.org maturity model (Crawl -> Walk -> Run), identifying gaps and opportunities across tagging, savings, and cost governance Tagging & Allocation Gaps - Identify and remediate gaps in cost allocation, tagging, and labelling across on-premises assets to enable accurate chargeback, showback, and cost attribution Savings Identification - Discover potential savings opportunities through licence optimisation, hardware rightsizing, capacity rebalancing, decommissioning, and vendor renegotiation Cost Governance Gaps - Assess and strengthen cost governance processes including approval workflows, budget controls, anomaly detection, and policy enforcement Data Normalisation - Normalise cost data from disparate sources into a unified cost model, enabling consistent reporting across business units, environments, and service categories Usage, Pattern & Licence Analysis - Analyse infrastructure utilisation patterns, workload distribution, peak/off-peak trends, and software licence usage to identify waste and optimisation opportunities Cost Modelling & Economics - Define unit economics, cost-per-service models, TCO calculations, and amortisation schedules for on-premises infrastructure to enable informed decision-making Implementation Roadmap - Develop a prioritised roadmap of quick-wins (0-3 months), mid-term initiatives (3-6 months), and long-term transformations (6-12 months) based on value and feasibility Reporting & Dashboards - Design and recommend live/near-real-time cost dashboards with drill-down capabilities, trend analysis, anomaly alerts, and executive reporting views FINOPS PERSONA COLLABORATION As defined by FinOps.org, this role works with all key FinOps personas: Executives (CTO, CFO, CIO) - Present cost visibility dashboards, business cases for optimisation, and FinOps maturity roadmaps to drive strategic investment decisions Finance / Procurement - Collaborate on cost allocation models, chargeback/showback mechanisms, budgeting, forecasting, and vendor contract analysis for on-premises licensing Engineering / DevOps Teams - Partner on tagging strategies, resource utilisation analysis, rightsizing recommendations, and embedding cost-awareness into engineering culture Product Owners / Business Units - Provide unit economics and cost-per-service metrics to inform product decisions, capacity planning, and build-vs-buy analysis IT Operations / Infrastructure - Work on data source integration, asset management, capacity utilisation, and hardware lifecycle cost modelling Architecture & Governance - Align FinOps governance policies, cost guardrails, and architectural decisions with total cost of ownership (TCO) analysis Security & Compliance - Ensure cost data handling meets compliance requirements and that FinOps tooling adheres to organisational security policies KEY RESPONSIBILITIES Define and own the on-premises FinOps strategy, architecture, and maturity roadmap aligned to the FinOps.org Inform -> Optimise -> Operate lifecycle Lead the FinOps maturity assessment (Crawl/Walk/Run) across tagging, savings identification, and cost governance - produce a baseline scorecard with prioritised improvement actions Architect the cost data model: identify and integrate all data sources (CMDB, asset management, procurement, licensing, monitoring), define normalisation rules, and establish the single source of truth for on-premises cost data Design cost allocation frameworks: define chargeback/showback models, tagging taxonomies, and cost attribution rules that enable accurate cost-per-business-unit, cost-per-service, and cost-per-environment reporting Develop unit economics and cost modelling: build TCO models, amortisation schedules, build-vs-buy comparisons, and cost-per-transaction metrics for on-premises infrastructure Lead savings discovery workshops across all FinOps personas: licence optimisation (true-ups, shelfware, harvesting), hardware rightsizing, capacity rebalancing, and decommissioning opportunities Architect the FinOps reporting and dashboard ecosystem: define requirements for live/near-real-time dashboards with executive views, drill-down capabilities, trend analysis, anomaly detection, and budget variance alerts Establish FinOps governance framework: cost approval policies, budget guardrails, anomaly response procedures, and regular cost review cadences with all persona groups Create the prioritised FinOps implementation roadmap: quick-wins (0-3 months), mid-term initiatives (3-6 months), and long-term transformation (6-12 months) based on value, feasibility, and stakeholder alignment Mentor and lead the FinOps engineering team (1 onshore + 1 offshore), establish engineering standards, and drive a culture of cost transparency and accountability across the organisation Present FinOps insights, savings opportunities, and governance reports to executive stakeholders (CTO, CFO, CIO) and lead regular FinOps review forums Stay current with FinOps.org framework evolution, contribute to the FinOps community of practice, and continuously refine methodologies for on-premises environments TECHNICAL SKILLS & EXPERTISE Deep expertise in on-premises infrastructure economics: server, storage, network, licensing (Microsoft EA, Oracle, VMware, Red Hat), data centre costs (power, cooling, space), and depreciation/amortisation models Strong understanding of the FinOps.org framework: Inform, Optimise, Operate lifecycle; FinOps Domains (Cost Allocation, Anomaly Management, Rate Optimisation, Usage Optimisation, etc.); FinOps maturity model (Crawl/Walk/Run) Experience with cost data integration from: CMDB (ServiceNow), asset management tools, procurement platforms, licence management tools (Snow, Flexera), and infrastructure monitoring (Datadog, Dynatrace, SolarWinds) Data modelling and analytics: SQL, Python/Pandas for cost data analysis; experience building normalisation pipelines and cost allocation engines Dashboard and reporting tools: Power BI, Tableau, Grafana, or custom BI solutions for real-time cost visibility and executive dashboards Understanding of ITIL asset and configuration management, hardware lifecycle management, and capacity planning methodologies Familiarity with FinOps tooling: Apptio, CloudHealth/VMware Aria, Flexera One, ServiceNow ITOM/ITFM, or custom-built cost management platforms Knowledge of hybrid/multi-cloud cost management to bridge on-premises FinOps with cloud FinOps practices where applicable SOFT SKILLS & COMPETENCIES Exceptional stakeholder management - able to influence executives, finance, engineering, and operations teams with equal credibility Strategic thinker who translates complex cost data into clear business narratives and actionable recommendations Strong leadership and mentoring ability - coaches and develops FinOps engineers and builds a cost-aware engineering culture Excellent presentation and communication skills - comfortable presenting to C-level audiences and leading governance forums Commercial acumen - understands P&L, depreciation, amortisation, OpEx/CapEx distinctions, and vendor contract structures . click apply for full job details
We're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people-centric and here to power good.Every day, we future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our companyand customers from what's now to what's next.We make it happen through the power of acceleration. Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us. Job Description Mandatory Skills: Observability, Resiliency, Service Management, Reliability, Performance engineering, Scalability, release management, Cloud cost management. Role Description Skills: ROLE PURPOSE Lead the Site Reliability Engineering practice, driving the transformation from reactive operations to proactive, engineering-led reliability. Own the definition and enforcement of non-functional requirements (NFRs) using FMEA-based resiliency frameworks, and champion observability, self-healing automation, automated incident management, and database operations automation. Ensure systems are resilient, performant, cost-optimised, and continuously improving. KEY RESPONSIBILITIES Define and enforce non-functional requirements (NFRs) for performance, scalability, availability, fault tolerance, and cost efficiency using FMEA-based failure analysis Design and implement self-healing automation for known failure patterns, reducing human intervention and on-call burden by 50%+ Build comprehensive observability stacks (metrics, logs, traces) with ML-driven anomaly detection and AIOps capabilities Lead automated incident management: detection, triage, escalation, remediation, and post-incident review automation Drive DB automation: automated provisioning, release management (UK focus), backup/restore, and operational request workflows for all database operations Define and track SLIs, SLOs, and error budgets across all critical services, using them to balance reliability with feature velocity Conduct chaos engineering exercises and game days to validate resiliency and uncover hidden failure modes Mentor 2 SRE Engineers, establish engineering standards, and build a culture of reliability and continuous improvement Collaborate with Platform Engineering and Cloud teams to embed reliability into infrastructure and deployment pipelines TECHNICAL SKILLS & EXPERTISE Expert-level observability: Prometheus, Grafana, ELK/OpenSearch, Jaeger/Zipkin, Datadog, or Dynatrace Strong experience with AIOps and ML-driven monitoring: PagerDuty, Moogsoft, BigPanda, or custom ML pipelines Deep knowledge of FMEA, fault tree analysis, and chaos engineering tools (Gremlin, LitmusChaos, Chaos Monkey) Database automation: strong SQL skills plus experience with automated DB provisioning, migration tools (Liquibase, Flyway), and DB release pipelines Proficiency in automation and scripting: Python, Go, Bash, with experience building self-healing runbooks Infrastructure knowledge: Kubernetes, cloud platforms (AWS/Azure/GCP), networking, and storage systems CI/CD and release engineering: Jenkins, GitLab CI, Spinnaker, ArgoCD for integrated DB and application releases Cost management: experience with FinOps principles, resource optimisation, and cloud spend analysis SOFT SKILLS & COMPETENCIES Strong leadership and mentoring ability - coaches and develops junior engineers Excellent stakeholder management and communication skills across all levels Strategic thinker who balances technical depth with business outcomes Proven ability to drive change, influence without authority, and build consensus Strong analytical and problem-solving mindset with attention to detail Ability to manage competing priorities across multiple workstreams simultaneously QUALIFICATIONS & EXPERIENCE 7+ years in SRE, DevOps, or production engineering with 3+ years in a senior or lead capacity Proven track record of improving availability, reducing MTTR, and implementing self-healing at scale Experience managing or automating database operations in enterprise environments Relevant certifications preferred: CKA, AWS DevOps Professional, Azure DevOps Expert, SRE Foundation Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience) DESIRABLE / NICE TO HAVE Published work or conference talks on SRE, observability, or chaos engineering Experience with service mesh (Istio) and distributed tracing at scale Background in financial services or regulated industry SRE practices About Us We're a global, team of innovators. Together, we harness engineering excellence and passion to co-create meaningful solutions to complex challenges. We turn organizations into data-driven leaders that can make a positive impact on their industries and society. If you believe that innovation can bring a better tomorrow closer to today, this is the place for you. Fostering innovation through diverse perspectives Hitachi is a global company operating across a wide range of industries and regions. One of the things that sets Hitachi apart is the diversity of our business and people, which drives our innovation and growth. We are committed to building an inclusive culture based on mutual respect and merit-based systems. We believe that when people feel valued, heard, and safe to express themselves, they do their best work. How we look after you We help take care of your today and tomorrow with industry-leading benefits, support, and services that look after your holistic health and wellbeing. We're also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We're always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you'll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with. We're proud to say we're an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
27/07/2026
Full time
We're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people-centric and here to power good.Every day, we future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our companyand customers from what's now to what's next.We make it happen through the power of acceleration. Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us. Job Description Mandatory Skills: Observability, Resiliency, Service Management, Reliability, Performance engineering, Scalability, release management, Cloud cost management. Role Description Skills: ROLE PURPOSE Lead the Site Reliability Engineering practice, driving the transformation from reactive operations to proactive, engineering-led reliability. Own the definition and enforcement of non-functional requirements (NFRs) using FMEA-based resiliency frameworks, and champion observability, self-healing automation, automated incident management, and database operations automation. Ensure systems are resilient, performant, cost-optimised, and continuously improving. KEY RESPONSIBILITIES Define and enforce non-functional requirements (NFRs) for performance, scalability, availability, fault tolerance, and cost efficiency using FMEA-based failure analysis Design and implement self-healing automation for known failure patterns, reducing human intervention and on-call burden by 50%+ Build comprehensive observability stacks (metrics, logs, traces) with ML-driven anomaly detection and AIOps capabilities Lead automated incident management: detection, triage, escalation, remediation, and post-incident review automation Drive DB automation: automated provisioning, release management (UK focus), backup/restore, and operational request workflows for all database operations Define and track SLIs, SLOs, and error budgets across all critical services, using them to balance reliability with feature velocity Conduct chaos engineering exercises and game days to validate resiliency and uncover hidden failure modes Mentor 2 SRE Engineers, establish engineering standards, and build a culture of reliability and continuous improvement Collaborate with Platform Engineering and Cloud teams to embed reliability into infrastructure and deployment pipelines TECHNICAL SKILLS & EXPERTISE Expert-level observability: Prometheus, Grafana, ELK/OpenSearch, Jaeger/Zipkin, Datadog, or Dynatrace Strong experience with AIOps and ML-driven monitoring: PagerDuty, Moogsoft, BigPanda, or custom ML pipelines Deep knowledge of FMEA, fault tree analysis, and chaos engineering tools (Gremlin, LitmusChaos, Chaos Monkey) Database automation: strong SQL skills plus experience with automated DB provisioning, migration tools (Liquibase, Flyway), and DB release pipelines Proficiency in automation and scripting: Python, Go, Bash, with experience building self-healing runbooks Infrastructure knowledge: Kubernetes, cloud platforms (AWS/Azure/GCP), networking, and storage systems CI/CD and release engineering: Jenkins, GitLab CI, Spinnaker, ArgoCD for integrated DB and application releases Cost management: experience with FinOps principles, resource optimisation, and cloud spend analysis SOFT SKILLS & COMPETENCIES Strong leadership and mentoring ability - coaches and develops junior engineers Excellent stakeholder management and communication skills across all levels Strategic thinker who balances technical depth with business outcomes Proven ability to drive change, influence without authority, and build consensus Strong analytical and problem-solving mindset with attention to detail Ability to manage competing priorities across multiple workstreams simultaneously QUALIFICATIONS & EXPERIENCE 7+ years in SRE, DevOps, or production engineering with 3+ years in a senior or lead capacity Proven track record of improving availability, reducing MTTR, and implementing self-healing at scale Experience managing or automating database operations in enterprise environments Relevant certifications preferred: CKA, AWS DevOps Professional, Azure DevOps Expert, SRE Foundation Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience) DESIRABLE / NICE TO HAVE Published work or conference talks on SRE, observability, or chaos engineering Experience with service mesh (Istio) and distributed tracing at scale Background in financial services or regulated industry SRE practices About Us We're a global, team of innovators. Together, we harness engineering excellence and passion to co-create meaningful solutions to complex challenges. We turn organizations into data-driven leaders that can make a positive impact on their industries and society. If you believe that innovation can bring a better tomorrow closer to today, this is the place for you. Fostering innovation through diverse perspectives Hitachi is a global company operating across a wide range of industries and regions. One of the things that sets Hitachi apart is the diversity of our business and people, which drives our innovation and growth. We are committed to building an inclusive culture based on mutual respect and merit-based systems. We believe that when people feel valued, heard, and safe to express themselves, they do their best work. How we look after you We help take care of your today and tomorrow with industry-leading benefits, support, and services that look after your holistic health and wellbeing. We're also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We're always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you'll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with. We're proud to say we're an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
Job Title Senior Platform Engineer Job Description So, who are we? IG is a FTSE 100 fintech operating across five continents, serving over 1.3m customers and handling billions of dollars in transactions - built on scale, trust, and proof. We didn't pivot to innovation; it's how we've always operated. What that means for the people who work here is real: genuinely complex problems to solve, the technology and resources to tackle them properly, and the kind of scope that's rare in established businesses. The bar is high - bring a curious and forward-thinking mindset and we'll give you the platform to define what comes next. Join us at IG - the future gets built here. Your team This role is part of our Platform Engineering team they are responsible for building the cloud and containers platform landscape to support the business applications via self-service capabilities and automation. Your role in the Team's Success The Senior Platform Engineer is a hands on technical contributor responsible for building, evolving, and maintaining cloud platform infrastructure to ensure it is secure, scalable, reliable, and cost effective. The Senior Platform Engineer embeds best practices into platform design and delivery across Engineering, Security and FinOps. They operate as a hands on expert, contributing to complex technical decisions, reducing platform risk, and enabling engineering teams to build and operate services with confidence, speed, and financial discipline. The role includes mentoring junior engineers and supporting knowledge sharing across the team. What you'll do Design and architect scalable AWS cloud infrastructure, developing IaC solutions using Terraform and leading cloud technology evaluation and standardisation initiatives Build, optimise, and maintain CI/CD pipelines for infrastructure deployment, establishing DevOps and DevSecOps best practices across the organisation Architect and manage Kubernetes environments (EKS, AKS, GKE) and implement platform engineering tools to deliver self service capabilities via an IDP (e.g. Backstage) Define and enforce monitoring, alerting, and observability standards to ensure infrastructure reliability and performance Foster collaboration between development and operations teams, bridging gaps to drive a unified engineering culture Mentor junior engineers, share knowledge across the team, and write and maintain documentation to support continuous learning and engineering best practices What you'll need for this role Cloud Infrastructure & Experience: 4+ years in cloud infrastructure engineering with hands on experience designing and administering complex platforms across AWS, GCP, and Azure, including deep knowledge of networking (VPC, Direct Connect, Transit Gateways, BGP, OSPF) and IaaS/PaaS architectural patterns Infrastructure-as-Code & Automation: 2+ years with IaC tools (Terraform or Ansible) and hands on experience with orchestration and automation platforms (AWX, Ansible Tower, Terragrunt), scripting in Python, Bash, or Go DevOps & DevSecOps: 2+ years designing and implementing CI/CD pipelines using Jenkins, GitLab, GitHub Actions, or equivalent, with strong knowledge of security best practices including secrets management, vulnerability scanning, and compliance Kubernetes & Containers: 2+ years of Kubernetes administration across on premises and cloud environments (EKS preferred, AKS, GKE), with proficiency in networking, RBAC, pod security, service mesh (Istio), and Docker image management Platform Engineering & Developer Tooling: Experience implementing developer portals (Backstage), artifact repositories (JFrog Artifactory, Nexus), API gateways, and self service platform engineering tools to support internal teams Observability & Reliability: Hands on experience with monitoring and observability platforms (Honeycomb, Prometheus, Grafana), including defining alerting standards and ensuring infrastructure reliability at scale Architecture & Design: Strong grasp of microservices architecture, cloud networking protocols (Layer 2/3, MPLS), DNS, load balancing, and traffic management, with the ability to contribute to and implement technical direction Version Control & Collaboration: Proficiency across version control systems (Git, GitHub, GitLab, Bitbucket) with the ability to operate autonomously and contribute to engineering standards across teams How we work We try to take a thoughtful approach to our ways of working as a company. We follow a hybrid working model with 3 days in the office - which we think balances the need to collaborate effectively and connect with each other. When it comes to how we deliver, there are 5 things we want everyone to do to drive high performance, better learning and career satisfaction: Lead and Inspire: Drives trust, alignment, and enthusiasm Think Big: Focus on the problems that most impact commercial outcomes Champion the client: Understand and prioritise client's needs Deliver at pace: Push for fast, sustainable growth Raise the bar: Take ownership, be accountable and share feedback We believe that diversity is vital to success, it fuels creativity, drives innovation and sets us up for global success. We're committed to building teams with a variety of perspectives and skills to help us realise our vision and strategy, that's why we encourage applications from people with diverse backgrounds and experiences to join us on this journey. Learn more about our D&I approach here. The Perks Thrive with tailored development programs, mentoring opportunities with leaders, and clear career progression. Expand your network through committees, sports and social clubs. Enjoy extra time off for volunteering and community work. Learn more about the Perks here! Join us for this exciting journey.
27/07/2026
Full time
Job Title Senior Platform Engineer Job Description So, who are we? IG is a FTSE 100 fintech operating across five continents, serving over 1.3m customers and handling billions of dollars in transactions - built on scale, trust, and proof. We didn't pivot to innovation; it's how we've always operated. What that means for the people who work here is real: genuinely complex problems to solve, the technology and resources to tackle them properly, and the kind of scope that's rare in established businesses. The bar is high - bring a curious and forward-thinking mindset and we'll give you the platform to define what comes next. Join us at IG - the future gets built here. Your team This role is part of our Platform Engineering team they are responsible for building the cloud and containers platform landscape to support the business applications via self-service capabilities and automation. Your role in the Team's Success The Senior Platform Engineer is a hands on technical contributor responsible for building, evolving, and maintaining cloud platform infrastructure to ensure it is secure, scalable, reliable, and cost effective. The Senior Platform Engineer embeds best practices into platform design and delivery across Engineering, Security and FinOps. They operate as a hands on expert, contributing to complex technical decisions, reducing platform risk, and enabling engineering teams to build and operate services with confidence, speed, and financial discipline. The role includes mentoring junior engineers and supporting knowledge sharing across the team. What you'll do Design and architect scalable AWS cloud infrastructure, developing IaC solutions using Terraform and leading cloud technology evaluation and standardisation initiatives Build, optimise, and maintain CI/CD pipelines for infrastructure deployment, establishing DevOps and DevSecOps best practices across the organisation Architect and manage Kubernetes environments (EKS, AKS, GKE) and implement platform engineering tools to deliver self service capabilities via an IDP (e.g. Backstage) Define and enforce monitoring, alerting, and observability standards to ensure infrastructure reliability and performance Foster collaboration between development and operations teams, bridging gaps to drive a unified engineering culture Mentor junior engineers, share knowledge across the team, and write and maintain documentation to support continuous learning and engineering best practices What you'll need for this role Cloud Infrastructure & Experience: 4+ years in cloud infrastructure engineering with hands on experience designing and administering complex platforms across AWS, GCP, and Azure, including deep knowledge of networking (VPC, Direct Connect, Transit Gateways, BGP, OSPF) and IaaS/PaaS architectural patterns Infrastructure-as-Code & Automation: 2+ years with IaC tools (Terraform or Ansible) and hands on experience with orchestration and automation platforms (AWX, Ansible Tower, Terragrunt), scripting in Python, Bash, or Go DevOps & DevSecOps: 2+ years designing and implementing CI/CD pipelines using Jenkins, GitLab, GitHub Actions, or equivalent, with strong knowledge of security best practices including secrets management, vulnerability scanning, and compliance Kubernetes & Containers: 2+ years of Kubernetes administration across on premises and cloud environments (EKS preferred, AKS, GKE), with proficiency in networking, RBAC, pod security, service mesh (Istio), and Docker image management Platform Engineering & Developer Tooling: Experience implementing developer portals (Backstage), artifact repositories (JFrog Artifactory, Nexus), API gateways, and self service platform engineering tools to support internal teams Observability & Reliability: Hands on experience with monitoring and observability platforms (Honeycomb, Prometheus, Grafana), including defining alerting standards and ensuring infrastructure reliability at scale Architecture & Design: Strong grasp of microservices architecture, cloud networking protocols (Layer 2/3, MPLS), DNS, load balancing, and traffic management, with the ability to contribute to and implement technical direction Version Control & Collaboration: Proficiency across version control systems (Git, GitHub, GitLab, Bitbucket) with the ability to operate autonomously and contribute to engineering standards across teams How we work We try to take a thoughtful approach to our ways of working as a company. We follow a hybrid working model with 3 days in the office - which we think balances the need to collaborate effectively and connect with each other. When it comes to how we deliver, there are 5 things we want everyone to do to drive high performance, better learning and career satisfaction: Lead and Inspire: Drives trust, alignment, and enthusiasm Think Big: Focus on the problems that most impact commercial outcomes Champion the client: Understand and prioritise client's needs Deliver at pace: Push for fast, sustainable growth Raise the bar: Take ownership, be accountable and share feedback We believe that diversity is vital to success, it fuels creativity, drives innovation and sets us up for global success. We're committed to building teams with a variety of perspectives and skills to help us realise our vision and strategy, that's why we encourage applications from people with diverse backgrounds and experiences to join us on this journey. Learn more about our D&I approach here. The Perks Thrive with tailored development programs, mentoring opportunities with leaders, and clear career progression. Expand your network through committees, sports and social clubs. Enjoy extra time off for volunteering and community work. Learn more about the Perks here! Join us for this exciting journey.
Senior Platform EngineerApplylocations: City of London - United Kingdomtime type: Full timeposted on: Posted Todayjob requisition id: R\_17523 Job Title Senior Platform Engineer Job Description # So, who are we? IG is a FTSE 100 fintech operating across five continents, serving over 1.3m customers and handling billions of dollars in transactions - built on scale, trust, and proof. We didn't pivot to innovation; it's how we've always operated.What that means for the people who work here is real: genuinely complex problems to solve, the technology and resources to tackle them properly, and the kind of scope that's rare in established businesses.The bar is high - bring a curious and forward-thinking mindset and we'll give you the platform to define what comes next. Join us at IG - the future gets built here.# Your team This role is part of our Platform Engineering team they are responsible for building the cloud and containers platform landscape to support the business applications via self-service capabilities and automation.# Your role in the Team's Success The Senior Platform Engineer is a hands-on technical contributor responsible for building, evolving, and maintaining cloud platform infrastructure to ensure it is secure, scalable, reliable, and cost-effective.The Senior Platform Engineer embeds best practices into platform design and delivery across Engineering, Security and FinOps. They operate as a hands-on expert, contributing to complex technical decisions, reducing platform risk, and enabling engineering teams to build and operate services with confidence, speed, and financial discipline. The role includes mentoring junior engineers and supporting knowledge sharing across the team.# What you'll do Design and architect scalable AWS cloud infrastructure, developing IaC solutions using Terraform and leading cloud technology evaluation and standardisation initiatives Build, optimise, and maintain CI/CD pipelines for infrastructure deployment, establishing DevOps and DevSecOps best practices across the organisation Architect and manage Kubernetes environments (EKS, AKS, GKE) and implement platform engineering tools to deliver self-service capabilities via an IDP (e.g. Backstage) Define and enforce monitoring, alerting, and observability standards to ensure infrastructure reliability and performance Foster collaboration between development and operations teams, bridging gaps to drive a unified engineering culture Mentor junior engineers, share knowledge across the team, and write and maintain documentation to support continuous learning and engineering best practices# What you'll need for this role Cloud Infrastructure & Experience : 4+ years in cloud infrastructure engineering with hands-on experience designing and administering complex platforms across AWS, GCP, and Azure, including deep knowledge of networking (VPC, Direct Connect, Transit Gateways, BGP, OSPF) and IaaS/PaaS architectural patterns Infrastructure-as-Code & Automation : 2+ years with IaC tools (Terraform or Ansible) and hands-on experience with orchestration and automation platforms (AWX, Ansible Tower, Terragrunt), scripting in Python, Bash, or Go DevOps & DevSecOps : 2+ years designing and implementing CI/CD pipelines using Jenkins, GitLab, GitHub Actions, or equivalent, with strong knowledge of security best practices including secrets management, vulnerability scanning, and compliance Kubernetes & Containers : 2+ years of Kubernetes administration across on-premises and cloud environments (EKS preferred, AKS, GKE), with proficiency in networking, RBAC, pod security, service mesh (Istio), and Docker image management Platform Engineering & Developer Tooling : Experience implementing developer portals (Backstage), artifact repositories (JFrog Artifactory, Nexus), API gateways, and self-service platform engineering tools to support internal teams Observability & Reliability : Hands-on experience with monitoring and observability platforms (Honeycomb, Prometheus, Grafana), including defining alerting standards and ensuring infrastructure reliability at scale Architecture & Design: Strong grasp of microservices architecture, cloud networking protocols (Layer 2/3, MPLS), DNS, load balancing, and traffic management, with the ability to contribute to and implement technical direction Version Control & Collaboration : Proficiency across version control systems (Git, GitHub, GitLab, Bitbucket) with the ability to operate autonomously and contribute to engineering standards across teams# How we work We try to take a thoughtful approach to our ways of working as a company. We follow a hybrid working model with 3 days in the office which we think balances the need to collaborate effectively and connect with each other. When it comes to how we deliver, there are 5 things we want everyone to do to drive high performance, better learning and career satisfaction: Lead and Inspire: Drives trust, alignment, and enthusiasm Think Big: Focus on the problems that most impact commercial outcomes Champion the client: Understand and prioritise client's needs Deliver at pace: Push for fast, sustainable growth; Raise the bar: Take ownership, be accountable and share feedbackWe believe that diversity is vital to success, it fuels creativity, drives innovation and sets us up for global success. We're committed to building teams with a variety of perspectives and skills to help us realise our vision and strategy, that's why we encourage applications from people with diverse backgrounds and experiences to join us on this journey. Learn more about our D&I approach here.# The Perks Your growth fuels our success! Thrive with tailored development programs, mentoring opportunities with leaders, and clear career progression. Expand your network through committees, sports and social clubs. Enjoy extra time off for volunteering and community work.Learn more about the Perks here! Join us for this exciting journey. Apply now!
27/07/2026
Full time
Senior Platform EngineerApplylocations: City of London - United Kingdomtime type: Full timeposted on: Posted Todayjob requisition id: R\_17523 Job Title Senior Platform Engineer Job Description # So, who are we? IG is a FTSE 100 fintech operating across five continents, serving over 1.3m customers and handling billions of dollars in transactions - built on scale, trust, and proof. We didn't pivot to innovation; it's how we've always operated.What that means for the people who work here is real: genuinely complex problems to solve, the technology and resources to tackle them properly, and the kind of scope that's rare in established businesses.The bar is high - bring a curious and forward-thinking mindset and we'll give you the platform to define what comes next. Join us at IG - the future gets built here.# Your team This role is part of our Platform Engineering team they are responsible for building the cloud and containers platform landscape to support the business applications via self-service capabilities and automation.# Your role in the Team's Success The Senior Platform Engineer is a hands-on technical contributor responsible for building, evolving, and maintaining cloud platform infrastructure to ensure it is secure, scalable, reliable, and cost-effective.The Senior Platform Engineer embeds best practices into platform design and delivery across Engineering, Security and FinOps. They operate as a hands-on expert, contributing to complex technical decisions, reducing platform risk, and enabling engineering teams to build and operate services with confidence, speed, and financial discipline. The role includes mentoring junior engineers and supporting knowledge sharing across the team.# What you'll do Design and architect scalable AWS cloud infrastructure, developing IaC solutions using Terraform and leading cloud technology evaluation and standardisation initiatives Build, optimise, and maintain CI/CD pipelines for infrastructure deployment, establishing DevOps and DevSecOps best practices across the organisation Architect and manage Kubernetes environments (EKS, AKS, GKE) and implement platform engineering tools to deliver self-service capabilities via an IDP (e.g. Backstage) Define and enforce monitoring, alerting, and observability standards to ensure infrastructure reliability and performance Foster collaboration between development and operations teams, bridging gaps to drive a unified engineering culture Mentor junior engineers, share knowledge across the team, and write and maintain documentation to support continuous learning and engineering best practices# What you'll need for this role Cloud Infrastructure & Experience : 4+ years in cloud infrastructure engineering with hands-on experience designing and administering complex platforms across AWS, GCP, and Azure, including deep knowledge of networking (VPC, Direct Connect, Transit Gateways, BGP, OSPF) and IaaS/PaaS architectural patterns Infrastructure-as-Code & Automation : 2+ years with IaC tools (Terraform or Ansible) and hands-on experience with orchestration and automation platforms (AWX, Ansible Tower, Terragrunt), scripting in Python, Bash, or Go DevOps & DevSecOps : 2+ years designing and implementing CI/CD pipelines using Jenkins, GitLab, GitHub Actions, or equivalent, with strong knowledge of security best practices including secrets management, vulnerability scanning, and compliance Kubernetes & Containers : 2+ years of Kubernetes administration across on-premises and cloud environments (EKS preferred, AKS, GKE), with proficiency in networking, RBAC, pod security, service mesh (Istio), and Docker image management Platform Engineering & Developer Tooling : Experience implementing developer portals (Backstage), artifact repositories (JFrog Artifactory, Nexus), API gateways, and self-service platform engineering tools to support internal teams Observability & Reliability : Hands-on experience with monitoring and observability platforms (Honeycomb, Prometheus, Grafana), including defining alerting standards and ensuring infrastructure reliability at scale Architecture & Design: Strong grasp of microservices architecture, cloud networking protocols (Layer 2/3, MPLS), DNS, load balancing, and traffic management, with the ability to contribute to and implement technical direction Version Control & Collaboration : Proficiency across version control systems (Git, GitHub, GitLab, Bitbucket) with the ability to operate autonomously and contribute to engineering standards across teams# How we work We try to take a thoughtful approach to our ways of working as a company. We follow a hybrid working model with 3 days in the office which we think balances the need to collaborate effectively and connect with each other. When it comes to how we deliver, there are 5 things we want everyone to do to drive high performance, better learning and career satisfaction: Lead and Inspire: Drives trust, alignment, and enthusiasm Think Big: Focus on the problems that most impact commercial outcomes Champion the client: Understand and prioritise client's needs Deliver at pace: Push for fast, sustainable growth; Raise the bar: Take ownership, be accountable and share feedbackWe believe that diversity is vital to success, it fuels creativity, drives innovation and sets us up for global success. We're committed to building teams with a variety of perspectives and skills to help us realise our vision and strategy, that's why we encourage applications from people with diverse backgrounds and experiences to join us on this journey. Learn more about our D&I approach here.# The Perks Your growth fuels our success! Thrive with tailored development programs, mentoring opportunities with leaders, and clear career progression. Expand your network through committees, sports and social clubs. Enjoy extra time off for volunteering and community work.Learn more about the Perks here! Join us for this exciting journey. Apply now!
FunctionCloud & Data EngineeringOur CompanyWe're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people-centric and here to power good. Every day, we future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our company and customers from what's now to what's next. We make it happen through the power of acceleration.Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us.Job descriptionMandatory Skills:Observability, Resiliency, Service Management, Reliability, Performance engineering, Scalability, release management, Cloud cost management.Role Description Skills:ROLE PURPOSELead the Site Reliability Engineering practice, driving the transformation from reactive operations to proactive, engineering-led reliability. Own the definition and enforcement of non-functional requirements (NFRs) using FMEA-based resiliency frameworks, and champion observability, self-healing automation, automated incident management, and database operations automation. Ensure systems are resilient, performant, cost-optimised, and continuously improving.KEY RESPONSIBILITIESDefine and enforce non-functional requirements (NFRs) for performance, scalability, availability, fault tolerance, and cost efficiency using FMEA-based failure analysisDesign and implement self-healing automation for known failure patterns, reducing human intervention and on-call burden by 50%+Build comprehensive observability stacks (metrics, logs, traces) with ML-driven anomaly detection and AIOps capabilitiesLead automated incident management: detection, triage, escalation, remediation, and post-incident review automationDrive DB automation: automated provisioning, release management (UK focus), backup/restore, and operational request workflows for all database operationsDefine and track SLIs, SLOs, and error budgets across all critical services, using them to balance reliability with feature velocityConduct chaos engineering exercises and game days to validate resiliency and uncover hidden failure modesMentor 2 SRE Engineers, establish engineering standards, and build a culture of reliability and continuous improvementCollaborate with Platform Engineering and Cloud teams to embed reliability into infrastructure and deployment pipelinesTECHNICAL SKILLS & EXPERTISEExpert-level observability: Prometheus, Grafana, ELK/OpenSearch, Jaeger/Zipkin, Datadog, or DynatraceStrong experience with AIOps and ML-driven monitoring: PagerDuty, Moogsoft, BigPanda, or custom ML pipelinesDeep knowledge of FMEA, fault tree analysis, and chaos engineering tools (Gremlin, LitmusChaos, Chaos Monkey)Database automation: strong SQL skills plus experience with automated DB provisioning, migration tools (Liquibase, Flyway), and DB release pipelinesProficiency in automation and scripting: Python, Go, Bash, with experience building self-healing runbooksInfrastructure knowledge: Kubernetes, cloud platforms (AWS/Azure/GCP), networking, and storage systemsCI/CD and release engineering: Jenkins, GitLab CI, Spinnaker, ArgoCD for integrated DB and application releasesCost management: experience with FinOps principles, resource optimisation, and cloud spend analysisSOFT SKILLS & COMPETENCIESStrong leadership and mentoring ability - coaches and develops junior engineersExcellent stakeholder management and communication skills across all levelsStrategic thinker who balances technical depth with business outcomesProven ability to drive change, influence without authority, and build consensusStrong analytical and problem-solving mindset with attention to detailAbility to manage competing priorities across multiple workstreams simultaneouslyQUALIFICATIONS & EXPERIENCE7+ years in SRE, DevOps, or production engineering with 3+ years in a senior or lead capacityProven track record of improving availability, reducing MTTR, and implementing self-healing at scaleExperience managing or automating database operations in enterprise environmentsRelevant certifications preferred: CKA, AWS DevOps Professional, Azure DevOps Expert, SRE FoundationBachelor's degree in Computer Science, Engineering, or related field (or equivalent experience)DESIRABLE / NICE TO HAVEPublished work or conference talks on SRE, observability, or chaos engineeringExperience with service mesh (Istio) and distributed tracing at scaleBackground in financial services or regulated industry SRE practicesAbout usWe're a global, team of innovators. Together, we harness engineering excellence and passion to co-create meaningful solutions to complex challenges. We turn organizations into data-driven leaders that can make a positive impact on their industries and society. If you believe that innovation can bring a better tomorrow closer to today, this is the place for you.Fostering innovation through diverse perspectivesHitachi is a global company operating across a wide range of industries and regions. One of the things that sets Hitachi apart is the diversity of our business and people, which drives our innovation and growth.We are committed to building an inclusive culture based on mutual respect and merit-based systems. We believe that when people feel valued, heard, and safe to express themselves, they do their best work.How we look after youWe help take care of your today and tomorrow with industry-leading benefits, support, and services that look after your holistic health and wellbeing. We're also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We're always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you'll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with.We're proud to say we're an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
26/07/2026
Full time
FunctionCloud & Data EngineeringOur CompanyWe're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people-centric and here to power good. Every day, we future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our company and customers from what's now to what's next. We make it happen through the power of acceleration.Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us.Job descriptionMandatory Skills:Observability, Resiliency, Service Management, Reliability, Performance engineering, Scalability, release management, Cloud cost management.Role Description Skills:ROLE PURPOSELead the Site Reliability Engineering practice, driving the transformation from reactive operations to proactive, engineering-led reliability. Own the definition and enforcement of non-functional requirements (NFRs) using FMEA-based resiliency frameworks, and champion observability, self-healing automation, automated incident management, and database operations automation. Ensure systems are resilient, performant, cost-optimised, and continuously improving.KEY RESPONSIBILITIESDefine and enforce non-functional requirements (NFRs) for performance, scalability, availability, fault tolerance, and cost efficiency using FMEA-based failure analysisDesign and implement self-healing automation for known failure patterns, reducing human intervention and on-call burden by 50%+Build comprehensive observability stacks (metrics, logs, traces) with ML-driven anomaly detection and AIOps capabilitiesLead automated incident management: detection, triage, escalation, remediation, and post-incident review automationDrive DB automation: automated provisioning, release management (UK focus), backup/restore, and operational request workflows for all database operationsDefine and track SLIs, SLOs, and error budgets across all critical services, using them to balance reliability with feature velocityConduct chaos engineering exercises and game days to validate resiliency and uncover hidden failure modesMentor 2 SRE Engineers, establish engineering standards, and build a culture of reliability and continuous improvementCollaborate with Platform Engineering and Cloud teams to embed reliability into infrastructure and deployment pipelinesTECHNICAL SKILLS & EXPERTISEExpert-level observability: Prometheus, Grafana, ELK/OpenSearch, Jaeger/Zipkin, Datadog, or DynatraceStrong experience with AIOps and ML-driven monitoring: PagerDuty, Moogsoft, BigPanda, or custom ML pipelinesDeep knowledge of FMEA, fault tree analysis, and chaos engineering tools (Gremlin, LitmusChaos, Chaos Monkey)Database automation: strong SQL skills plus experience with automated DB provisioning, migration tools (Liquibase, Flyway), and DB release pipelinesProficiency in automation and scripting: Python, Go, Bash, with experience building self-healing runbooksInfrastructure knowledge: Kubernetes, cloud platforms (AWS/Azure/GCP), networking, and storage systemsCI/CD and release engineering: Jenkins, GitLab CI, Spinnaker, ArgoCD for integrated DB and application releasesCost management: experience with FinOps principles, resource optimisation, and cloud spend analysisSOFT SKILLS & COMPETENCIESStrong leadership and mentoring ability - coaches and develops junior engineersExcellent stakeholder management and communication skills across all levelsStrategic thinker who balances technical depth with business outcomesProven ability to drive change, influence without authority, and build consensusStrong analytical and problem-solving mindset with attention to detailAbility to manage competing priorities across multiple workstreams simultaneouslyQUALIFICATIONS & EXPERIENCE7+ years in SRE, DevOps, or production engineering with 3+ years in a senior or lead capacityProven track record of improving availability, reducing MTTR, and implementing self-healing at scaleExperience managing or automating database operations in enterprise environmentsRelevant certifications preferred: CKA, AWS DevOps Professional, Azure DevOps Expert, SRE FoundationBachelor's degree in Computer Science, Engineering, or related field (or equivalent experience)DESIRABLE / NICE TO HAVEPublished work or conference talks on SRE, observability, or chaos engineeringExperience with service mesh (Istio) and distributed tracing at scaleBackground in financial services or regulated industry SRE practicesAbout usWe're a global, team of innovators. Together, we harness engineering excellence and passion to co-create meaningful solutions to complex challenges. We turn organizations into data-driven leaders that can make a positive impact on their industries and society. If you believe that innovation can bring a better tomorrow closer to today, this is the place for you.Fostering innovation through diverse perspectivesHitachi is a global company operating across a wide range of industries and regions. One of the things that sets Hitachi apart is the diversity of our business and people, which drives our innovation and growth.We are committed to building an inclusive culture based on mutual respect and merit-based systems. We believe that when people feel valued, heard, and safe to express themselves, they do their best work.How we look after youWe help take care of your today and tomorrow with industry-leading benefits, support, and services that look after your holistic health and wellbeing. We're also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We're always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you'll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with.We're proud to say we're an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.
We're looking for a Senior Data Platform Engineer to join our team on a 12-month fixed term contract, to play a pivotal role in shaping and operating our next-generation data platform. This is a fantastic opportunity to join our Microlise One Analytics & AI platform, where you'll help build a secure, scalable and governed data ecosystem that powers analytics, AI, and business critical insights across the organisation. You'll blend platform engineering, cloud DevOps, and data governance, ensuring our platform is not only high performing, but also compliant, reliable, and cost efficient. Be part of a high impact data and AI transformation programme, work on a modern, cloud native platform with real business impact, influence how data, analytics and AI are delivered across the organisation, and collaborate with cross functional engineering and data teams. What you'll be doing Designing, building and operating our AWS-based data platform using infrastructure as code Creating and managing CI/CD pipelines to enable safe, controlled deployments Supporting modern Lakehouse architecture, ingestion, processing and data serving layers Embedding security, governance and compliance controls across the platform Driving platform observability, monitoring, and performance optimisation Owning aspects of cost management (FinOps) and platform efficiency Collaborating with Data Engineers, Analytics & AI teams to deliver trusted data products What you'll bring Strong experience in platform engineering, cloud engineering or DevOps within data driven environments Hands on expertise in AWS, infrastructure as code and CI/CD pipelines Experience building secure, scalable cloud environments A solid understanding of data platforms, pipelines or analytics ecosystems A mindset focused on reliability, observability and continuous improvement Technical skills: AWS (compute, storage, IAM, networking) Terraform and/or AWS CDK CI/CD tooling and pipeline engineering Observability, monitoring and operational tooling SQL and strong systems thinking Experience in regulated or data sensitive environments Multi account or multi tenant cloud architecture exposure Why Microlise? When your groceries arrive at your door or you sign for your online parcel, one or more of our software, telematics or proof of purchase solutions has probably been used. Our solutions deliver value to many of the UK's leading grocery retailers and food logistics providers as well as to household names including JCB, Eddie Stobart, Carlsberg, Waitrose, and Royal Mail. Proudly Midlands based, Microlise has been operating for over thirty years, and recently became a Publicly Listed Company with shares trading on the London Stock Exchange. Our growing business is guided by our culture which drives the way we behave, the way we work, the way we connect with our customers, and the way we support and develop our people. We believe in developing our staff and support our employees with their professional development goals 37.5 hour week with flexible working opportunities Access to our salary sacrifice EV Car Scheme - payments are made before tax and other contributions, so saving you money, whilst doing your bit for the environment! Great Place to Work certified - We have been recognised by the global authority on workplace culture, so come be a part of our success! Private medical insurance with Vitality Health including rewards for members such as: Free Amazon Prime, Apple Watch, discounted gym membership and many more! 25 days holiday, excluding bank holidays, increasing with service Invested in employee health and well being with over 20 mental health first aiders in the business Employee Assistance Programmes Free Costco membership, 20% off EE mobile and line rental, and other local discounts Great staff extras: Easter eggs, yearly BBQ, Christmas gifts and annual staff awards Executive Box at Motorpoint Arena Nottingham Recruitment Process For successful candidates, interviews will take place whilst the advert is still live, via telephone and video conferencing; so don't delay getting your application in! Recruitment Agencies Whilst we make every effort to directly source candidates for our live roles, we do have a very small preferred supplier list on the occasion we may require additional support. We therefore do not accept speculative CVs and/or cold calls to our Recruitment Team or Hiring Managers. Any queries should be directed to in the first instance.
26/07/2026
Full time
We're looking for a Senior Data Platform Engineer to join our team on a 12-month fixed term contract, to play a pivotal role in shaping and operating our next-generation data platform. This is a fantastic opportunity to join our Microlise One Analytics & AI platform, where you'll help build a secure, scalable and governed data ecosystem that powers analytics, AI, and business critical insights across the organisation. You'll blend platform engineering, cloud DevOps, and data governance, ensuring our platform is not only high performing, but also compliant, reliable, and cost efficient. Be part of a high impact data and AI transformation programme, work on a modern, cloud native platform with real business impact, influence how data, analytics and AI are delivered across the organisation, and collaborate with cross functional engineering and data teams. What you'll be doing Designing, building and operating our AWS-based data platform using infrastructure as code Creating and managing CI/CD pipelines to enable safe, controlled deployments Supporting modern Lakehouse architecture, ingestion, processing and data serving layers Embedding security, governance and compliance controls across the platform Driving platform observability, monitoring, and performance optimisation Owning aspects of cost management (FinOps) and platform efficiency Collaborating with Data Engineers, Analytics & AI teams to deliver trusted data products What you'll bring Strong experience in platform engineering, cloud engineering or DevOps within data driven environments Hands on expertise in AWS, infrastructure as code and CI/CD pipelines Experience building secure, scalable cloud environments A solid understanding of data platforms, pipelines or analytics ecosystems A mindset focused on reliability, observability and continuous improvement Technical skills: AWS (compute, storage, IAM, networking) Terraform and/or AWS CDK CI/CD tooling and pipeline engineering Observability, monitoring and operational tooling SQL and strong systems thinking Experience in regulated or data sensitive environments Multi account or multi tenant cloud architecture exposure Why Microlise? When your groceries arrive at your door or you sign for your online parcel, one or more of our software, telematics or proof of purchase solutions has probably been used. Our solutions deliver value to many of the UK's leading grocery retailers and food logistics providers as well as to household names including JCB, Eddie Stobart, Carlsberg, Waitrose, and Royal Mail. Proudly Midlands based, Microlise has been operating for over thirty years, and recently became a Publicly Listed Company with shares trading on the London Stock Exchange. Our growing business is guided by our culture which drives the way we behave, the way we work, the way we connect with our customers, and the way we support and develop our people. We believe in developing our staff and support our employees with their professional development goals 37.5 hour week with flexible working opportunities Access to our salary sacrifice EV Car Scheme - payments are made before tax and other contributions, so saving you money, whilst doing your bit for the environment! Great Place to Work certified - We have been recognised by the global authority on workplace culture, so come be a part of our success! Private medical insurance with Vitality Health including rewards for members such as: Free Amazon Prime, Apple Watch, discounted gym membership and many more! 25 days holiday, excluding bank holidays, increasing with service Invested in employee health and well being with over 20 mental health first aiders in the business Employee Assistance Programmes Free Costco membership, 20% off EE mobile and line rental, and other local discounts Great staff extras: Easter eggs, yearly BBQ, Christmas gifts and annual staff awards Executive Box at Motorpoint Arena Nottingham Recruitment Process For successful candidates, interviews will take place whilst the advert is still live, via telephone and video conferencing; so don't delay getting your application in! Recruitment Agencies Whilst we make every effort to directly source candidates for our live roles, we do have a very small preferred supplier list on the occasion we may require additional support. We therefore do not accept speculative CVs and/or cold calls to our Recruitment Team or Hiring Managers. Any queries should be directed to in the first instance.
Goldman Sachs seeks a senior technical leader to join the Cloud Engineering & Architecture team in London, driving large-scale cloud initiatives, architecting resilient AWS infrastructure, and mentoring engineers across teams. We value experience designing autonomous, self-healing systems, implementing AI-driven FinOps, and aligning platform work with regulatory and governance requirements in a fast-paced financial services environment.
26/07/2026
Full time
Goldman Sachs seeks a senior technical leader to join the Cloud Engineering & Architecture team in London, driving large-scale cloud initiatives, architecting resilient AWS infrastructure, and mentoring engineers across teams. We value experience designing autonomous, self-healing systems, implementing AI-driven FinOps, and aligning platform work with regulatory and governance requirements in a fast-paced financial services environment.
What We Do At Goldman Sachs, our Engineers don't just make things - we make things possible. Change the world by connecting people and capital with ideas. Solve the most challenging and pressing engineering problems for our clients. Join our engineering teams that build massively scalable software and systems, architect low latency infrastructure solutions, proactively guard against cyber threats, and leverage machine learning alongside financial engineering to continuously turn data into action. Create new businesses, transform finance, and explore a world of opportunity at the speed of markets. Goldman Sachs Engineers are innovators and problem-solvers, building solutions in Artificial Intelligence, risk management, big data, mobile and more. As part of Core Engineering at Goldman Sachs, the CE&A team is responsible for enabling the use of public cloud services across the firm. You will be working as part of a multi-disciplinary team responsible for researching, architecting and building a cutting edge platform that enable Goldman Sachs Engineering teams to deploy and manage services in public cloud safely and securely. The organization is seeking highly collaborative, creative, and intellectually curious engineers who are passionate about developing and implementing cutting edge cloud computing and AI solutions. The ideal candidate will thrive in a DevOps culture and contribute to customer centric product development. They will work closely with cross functional teams, and will be creative collaborators who evolve, adapt to change and thrive in a fast paced global environment. Responsibilities And Qualifications: We are looking for a senior technical leader to join our Cloud Engineering & Architecture team and play a pivotal role in enabling the firm to maximize its use of cloud infrastructure. This is a hands on leadership position requiring deep technical expertise, strategic thinking, and the ability to drive large scale platform initiatives from conception to delivery. Key Responsibilities: Design, develop, and operationalize enterprise grade cloud platform capabilities. Architect scalable, resilient, and secure infrastructure solutions on AWS. Architect and operationalize autonomous AI based, self healing infrastructure Define technical standards, best practices, and reference architectures for cloud adoption across the firm. Partner with engineering teams to enable seamless migration and modernization of workloads to the cloud. Drive automation and infrastructure as code practices to improve operational efficiency. Drive AI powered FinOps and predictive resource optimization Mentor and guide engineers across teams, raising the overall technical bar. Collaborate with security, networking, and compliance teams to ensure platform meets regulatory and governance requirements. Evaluate emerging technologies and make recommendations for platform evolution. Participate in architecture design reviews and provide technical leadership on complex initiatives. Complementary AI Based Skills: LLM Orchestration & Agentic Workflows: Experience in designing, building, and deploying Large Language Model (LLM) orchestration frameworks (e.g., LangChain, Temporal, or custom agentic loops) to coordinate multi step diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root cause analysis, anomaly detection, and semantic log clustering. Self Healing Infrastructure Engineering: Experience designing closed loop, self healing systems that autonomously execute recovery actions (e.g., traffic shifting, automated rollbacks, or service restarts) with built in verification and safety guardrails. AI Driven FinOps & Resource Optimization: Deep understanding of applying machine learning and predictive analytics to dynamically right size cloud resources, manage spot instances, and optimize data platform workloads. Predictive Capacity Planning: Ability to design algorithms that forecast workload demands and proactively scale infrastructure to prevent over provisioning while maintaining strict SLAs. Complementary Behaviours: Toil Reduction Mindset: A relentless focus on eliminating repetitive operational support and engineering friction by shifting platform operations from reactive troubleshooting to autonomous mitigation. Risk Aware Automation: Demonstrates a disciplined approach to safety by implementing strict confidence thresholds, validation loops, and human in the loop fallbacks for autonomous AI actions. Value Oriented Engineering: Treats cost optimization as a first class architectural metric, aligning infrastructure spend directly with business value and platform efficiency. Impact Preserving Innovation: Executes large scale optimization initiatives with a meticulous, risk mitigated approach, ensuring zero disruption to production environments or developer velocity. Basic Qualifications: 10+ years of experience in software engineering or infrastructure engineering. Deep hands on expertise with AWS services (EC2, EKS, Lambda, S3, IAM, VPC, CloudFormation, CDK, etc.). Strong background in platform engineering, building internal developer platforms, or infrastructure tooling. Experience designing and operating large scale distributed systems. Proficiency with infrastructure as code tools (Terraform, CloudFormation). Strong understanding of containerization and orchestration (Docker, Kubernetes). Experience with CI/CD pipelines and DevOps practices. Knowledge of networking, security, and identity management in cloud environments. Excellent communication skills with the ability to influence technical decisions across teams. Experience working in regulated industries is a plus. Preferred Qualifications: Experience building self service platforms for development teams. Familiarity with observability and monitoring tools (Prometheus, Grafana, Datadog, CloudWatch). Background in financial services or other highly regulated environments. About Goldman Sachs At Goldman Sachs, we commit our people, capital and ideas to help our clients, shareholders and the communities we serve to grow. Founded in 1869, we are a leading global investment banking, securities and investment management firm. Headquartered in New York, we maintain offices around the world. We believe who you are makes you better at what you do. We're committed to fostering and advancing diversity and inclusion in our own workplace and beyond by ensuring every individual within our firm has several opportunities to grow professionally and personally, from our training and development opportunities and firmwide networks to benefits, wellness and personal finance offerings and mindfulness programs. Learn more about our culture, benefits, and people at We offer a wide range of health and welfare programs that vary depending on office location. These generally include medical, dental, short term disability, long term disability, life, accidental death, labor accident and business travel accident insurance. We offer competitive vacation policies based on employee level and office location. We promote time off from work to recharge by providing generous vacation entitlements and a minimum of three weeks expected vacation usage each year. Financial Wellness & Retirement We assist employees in saving and planning for retirement, offer financial support for higher education, and provide a number of benefits to help employees prepare for the unexpected. We offer live financial education and content on a variety of topics to address the spectrum of employees' priorities. Health Services We offer a medical advocacy service for employees and family members facing critical health situations, and counseling and referral services through the Employee Assistance Program (EAP). We provide Global Medical, Security and Travel Assistance and a Workplace Ergonomics Program. We also offer state of the art on site health centers in certain offices. Fitness To encourage employees to live a healthy and active lifestyle, some of our offices feature on site fitness centers. For eligible employees we typically reimburse fees paid for a fitness club membership or activity (up to a pre approved amount). Child Care & Family Care We offer on site child care centers that provide full time and emergency back up care, as well as mother and baby rooms and homework rooms. In every office, we provide advice and counseling services, expectant parent resources and transitional programs for parents returning from parental leave. Adoption, surrogacy, egg donation and egg retrieval stipends are also available. Benefits at Goldman Sachs Read more about the full suite of class leading benefits our firm has to offer.
26/07/2026
Full time
What We Do At Goldman Sachs, our Engineers don't just make things - we make things possible. Change the world by connecting people and capital with ideas. Solve the most challenging and pressing engineering problems for our clients. Join our engineering teams that build massively scalable software and systems, architect low latency infrastructure solutions, proactively guard against cyber threats, and leverage machine learning alongside financial engineering to continuously turn data into action. Create new businesses, transform finance, and explore a world of opportunity at the speed of markets. Goldman Sachs Engineers are innovators and problem-solvers, building solutions in Artificial Intelligence, risk management, big data, mobile and more. As part of Core Engineering at Goldman Sachs, the CE&A team is responsible for enabling the use of public cloud services across the firm. You will be working as part of a multi-disciplinary team responsible for researching, architecting and building a cutting edge platform that enable Goldman Sachs Engineering teams to deploy and manage services in public cloud safely and securely. The organization is seeking highly collaborative, creative, and intellectually curious engineers who are passionate about developing and implementing cutting edge cloud computing and AI solutions. The ideal candidate will thrive in a DevOps culture and contribute to customer centric product development. They will work closely with cross functional teams, and will be creative collaborators who evolve, adapt to change and thrive in a fast paced global environment. Responsibilities And Qualifications: We are looking for a senior technical leader to join our Cloud Engineering & Architecture team and play a pivotal role in enabling the firm to maximize its use of cloud infrastructure. This is a hands on leadership position requiring deep technical expertise, strategic thinking, and the ability to drive large scale platform initiatives from conception to delivery. Key Responsibilities: Design, develop, and operationalize enterprise grade cloud platform capabilities. Architect scalable, resilient, and secure infrastructure solutions on AWS. Architect and operationalize autonomous AI based, self healing infrastructure Define technical standards, best practices, and reference architectures for cloud adoption across the firm. Partner with engineering teams to enable seamless migration and modernization of workloads to the cloud. Drive automation and infrastructure as code practices to improve operational efficiency. Drive AI powered FinOps and predictive resource optimization Mentor and guide engineers across teams, raising the overall technical bar. Collaborate with security, networking, and compliance teams to ensure platform meets regulatory and governance requirements. Evaluate emerging technologies and make recommendations for platform evolution. Participate in architecture design reviews and provide technical leadership on complex initiatives. Complementary AI Based Skills: LLM Orchestration & Agentic Workflows: Experience in designing, building, and deploying Large Language Model (LLM) orchestration frameworks (e.g., LangChain, Temporal, or custom agentic loops) to coordinate multi step diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root cause analysis, anomaly detection, and semantic log clustering. Self Healing Infrastructure Engineering: Experience designing closed loop, self healing systems that autonomously execute recovery actions (e.g., traffic shifting, automated rollbacks, or service restarts) with built in verification and safety guardrails. AI Driven FinOps & Resource Optimization: Deep understanding of applying machine learning and predictive analytics to dynamically right size cloud resources, manage spot instances, and optimize data platform workloads. Predictive Capacity Planning: Ability to design algorithms that forecast workload demands and proactively scale infrastructure to prevent over provisioning while maintaining strict SLAs. Complementary Behaviours: Toil Reduction Mindset: A relentless focus on eliminating repetitive operational support and engineering friction by shifting platform operations from reactive troubleshooting to autonomous mitigation. Risk Aware Automation: Demonstrates a disciplined approach to safety by implementing strict confidence thresholds, validation loops, and human in the loop fallbacks for autonomous AI actions. Value Oriented Engineering: Treats cost optimization as a first class architectural metric, aligning infrastructure spend directly with business value and platform efficiency. Impact Preserving Innovation: Executes large scale optimization initiatives with a meticulous, risk mitigated approach, ensuring zero disruption to production environments or developer velocity. Basic Qualifications: 10+ years of experience in software engineering or infrastructure engineering. Deep hands on expertise with AWS services (EC2, EKS, Lambda, S3, IAM, VPC, CloudFormation, CDK, etc.). Strong background in platform engineering, building internal developer platforms, or infrastructure tooling. Experience designing and operating large scale distributed systems. Proficiency with infrastructure as code tools (Terraform, CloudFormation). Strong understanding of containerization and orchestration (Docker, Kubernetes). Experience with CI/CD pipelines and DevOps practices. Knowledge of networking, security, and identity management in cloud environments. Excellent communication skills with the ability to influence technical decisions across teams. Experience working in regulated industries is a plus. Preferred Qualifications: Experience building self service platforms for development teams. Familiarity with observability and monitoring tools (Prometheus, Grafana, Datadog, CloudWatch). Background in financial services or other highly regulated environments. About Goldman Sachs At Goldman Sachs, we commit our people, capital and ideas to help our clients, shareholders and the communities we serve to grow. Founded in 1869, we are a leading global investment banking, securities and investment management firm. Headquartered in New York, we maintain offices around the world. We believe who you are makes you better at what you do. We're committed to fostering and advancing diversity and inclusion in our own workplace and beyond by ensuring every individual within our firm has several opportunities to grow professionally and personally, from our training and development opportunities and firmwide networks to benefits, wellness and personal finance offerings and mindfulness programs. Learn more about our culture, benefits, and people at We offer a wide range of health and welfare programs that vary depending on office location. These generally include medical, dental, short term disability, long term disability, life, accidental death, labor accident and business travel accident insurance. We offer competitive vacation policies based on employee level and office location. We promote time off from work to recharge by providing generous vacation entitlements and a minimum of three weeks expected vacation usage each year. Financial Wellness & Retirement We assist employees in saving and planning for retirement, offer financial support for higher education, and provide a number of benefits to help employees prepare for the unexpected. We offer live financial education and content on a variety of topics to address the spectrum of employees' priorities. Health Services We offer a medical advocacy service for employees and family members facing critical health situations, and counseling and referral services through the Employee Assistance Program (EAP). We provide Global Medical, Security and Travel Assistance and a Workplace Ergonomics Program. We also offer state of the art on site health centers in certain offices. Fitness To encourage employees to live a healthy and active lifestyle, some of our offices feature on site fitness centers. For eligible employees we typically reimburse fees paid for a fitness club membership or activity (up to a pre approved amount). Child Care & Family Care We offer on site child care centers that provide full time and emergency back up care, as well as mother and baby rooms and homework rooms. In every office, we provide advice and counseling services, expectant parent resources and transitional programs for parents returning from parental leave. Adoption, surrogacy, egg donation and egg retrieval stipends are also available. Benefits at Goldman Sachs Read more about the full suite of class leading benefits our firm has to offer.
What We Do At Goldman Sachs, our Engineers don't just make things - we make things possible. Change the world by connecting people and capital with ideas. Solve the most challenging and pressing engineering problems for our clients. Join our engineering teams that build massively scalable software and systems, architect low latency infrastructure solutions, proactively guard against cyber threats, and leverage machine learning alongside financial engineering to continuously turn data into action. Create new businesses, transform finance, and explore a world of opportunity at the speed of markets. Goldman Sachs Engineers are innovators and problem-solvers, building solutions in Artificial Intelligence, risk management, big data, mobile and more. Cloud Engineering & Architecture (CE&A) As part of Core Engineering at Goldman Sachs, the CE&A team is responsible for enabling the use of public cloud services across the firm. You will be working as part of a multi-disciplinary team responsible for researching, architecting and building a cutting-edge platform that enable Goldman Sachs Engineering teams to deploy and manage services in public cloud safely and securely. The organization is seeking highly collaborative, creative, and intellectually curious engineers who are passionate about developing and implementing cutting-edge cloud computing and AI solutions. The ideal candidate will thrive in a DevOps culture and contribute to customer-centric product development. They will work closely with cross-functional teams, and will be creative collaborators who evolve, adapt to change and thrive in a fast-paced global environment. Responsibilities And Qualifications: We are looking for a senior technical leader to join our Cloud Engineering & Architecture team and play a pivotal role in enabling the firm to maximize its use of cloud infrastructure. This is a hands-on leadership position requiring deep technical expertise, strategic thinking, and the ability to drive large-scale platform initiatives from conception to delivery. Key Responsibilities: Design, develop, and operationalize enterprise-grade cloud platform capabilities. Architect scalable, resilient, and secure infrastructure solutions on AWS. Architect and operationalize autonomous AI based, self-healing infrastructure Define technical standards, best practices, and reference architectures for cloud adoption across the firm. Partner with engineering teams to enable seamless migration and modernization of workloads to the cloud. Drive automation and infrastructure-as-code practices to improve operational efficiency. Drive AI-powered FinOps and predictive resource optimization Mentor and guide engineers across teams, raising the overall technical bar. Collaborate with security, networking, and compliance teams to ensure platform meets regulatory and governance requirements. Evaluate emerging technologies and make recommendations for platform evolution. Participate in architecture design reviews and provide technical leadership on complex initiatives. Complementary AI Based Skills: LLM Orchestration & Agentic Workflows: Experience in designing, building, and deploying Large Language Model (LLM) orchestration frameworks (e.g., LangChain, Temporal, or custom agentic loops) to coordinate multi-step diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root-cause analysis, anomaly detection, and semantic log clustering. Self-Healing Infrastructure Engineering: Experience designing closed-loop, self-healing systems that autonomously execute recovery actions (e.g., traffic shifting, automated rollbacks, or service restarts) with built-in verification and safety guardrails. AI-Driven FinOps & Resource Optimization: Deep understanding of applying machine learning and predictive analytics to dynamically right-size cloud resources, manage spot instances, and optimize data platform workloads. Predictive Capacity Planning: Ability to design algorithms that forecast workload demands and proactively scale infrastructure to prevent over-provisioning while maintaining strict SLAs. Complementary Behaviours: Toil-Reduction Mindset: A relentless focus on eliminating repetitive operational support and engineering friction by shifting platform operations from reactive troubleshooting to autonomous mitigation. Risk-Aware Automation: Demonstrates a disciplined approach to safety by implementing strict confidence thresholds, validation loops, and human-in-the-loop fallbacks for autonomous AI actions. Value-Oriented Engineering: Treats cost optimization as a first-class architectural metric, aligning infrastructure spend directly with business value and platform efficiency. Impact-Preserving Innovation: Executes large-scale optimization initiatives with a meticulous, risk-mitigated approach, ensuring zero disruption to production environments or developer velocity. Basic Qualifications: 10+ years of experience in software engineering or infrastructure engineering. Deep hands-on expertise with AWS services (EC2, EKS, Lambda, S3, IAM, VPC, CloudFormation, CDK, etc.). Strong background in platform engineering, building internal developer platforms, or infrastructure tooling. Experience designing and operating large-scale distributed systems. Proficiency with infrastructure-as-code tools (Terraform, CloudFormation). Strong understanding of containerization and orchestration (Docker, Kubernetes). Experience with CI/CD pipelines and DevOps practices. Knowledge of networking, security, and identity management in cloud environments. Excellent communication skills with the ability to influence technical decisions across teams. Experience working in regulated industries is a plus. Preferred Qualifications: AWS certifications (Solutions Architect Professional, DevOps Engineer, etc.). Experience building self-service platforms for development teams. Familiarity with observability and monitoring tools (Prometheus, Grafana, Datadog, CloudWatch). Background in financial services or other highly regulated environments. About Goldman Sachs At Goldman Sachs, we commit our people, capital and ideas to help our clients, shareholders and the communities we serve to grow. Founded in 1869, we are a leading global investment banking, securities and investment management firm. Headquartered in New York, we maintain offices around the world. We believe who you are makes you better at what you do. We're committed to fostering and advancing diversity and inclusion in our own workplace and beyond by ensuring every individual within our firm has several opportunities to grow professionally and personally, from our training and development opportunities and firmwide networks to benefits, wellness and personal finance offerings and mindfulness programs. Learn more about our culture, benefits, and people at . We're committed to finding reasonable accommodation for candidates with special needs or disabilities during our recruiting process. Learn more:
26/07/2026
Full time
What We Do At Goldman Sachs, our Engineers don't just make things - we make things possible. Change the world by connecting people and capital with ideas. Solve the most challenging and pressing engineering problems for our clients. Join our engineering teams that build massively scalable software and systems, architect low latency infrastructure solutions, proactively guard against cyber threats, and leverage machine learning alongside financial engineering to continuously turn data into action. Create new businesses, transform finance, and explore a world of opportunity at the speed of markets. Goldman Sachs Engineers are innovators and problem-solvers, building solutions in Artificial Intelligence, risk management, big data, mobile and more. Cloud Engineering & Architecture (CE&A) As part of Core Engineering at Goldman Sachs, the CE&A team is responsible for enabling the use of public cloud services across the firm. You will be working as part of a multi-disciplinary team responsible for researching, architecting and building a cutting-edge platform that enable Goldman Sachs Engineering teams to deploy and manage services in public cloud safely and securely. The organization is seeking highly collaborative, creative, and intellectually curious engineers who are passionate about developing and implementing cutting-edge cloud computing and AI solutions. The ideal candidate will thrive in a DevOps culture and contribute to customer-centric product development. They will work closely with cross-functional teams, and will be creative collaborators who evolve, adapt to change and thrive in a fast-paced global environment. Responsibilities And Qualifications: We are looking for a senior technical leader to join our Cloud Engineering & Architecture team and play a pivotal role in enabling the firm to maximize its use of cloud infrastructure. This is a hands-on leadership position requiring deep technical expertise, strategic thinking, and the ability to drive large-scale platform initiatives from conception to delivery. Key Responsibilities: Design, develop, and operationalize enterprise-grade cloud platform capabilities. Architect scalable, resilient, and secure infrastructure solutions on AWS. Architect and operationalize autonomous AI based, self-healing infrastructure Define technical standards, best practices, and reference architectures for cloud adoption across the firm. Partner with engineering teams to enable seamless migration and modernization of workloads to the cloud. Drive automation and infrastructure-as-code practices to improve operational efficiency. Drive AI-powered FinOps and predictive resource optimization Mentor and guide engineers across teams, raising the overall technical bar. Collaborate with security, networking, and compliance teams to ensure platform meets regulatory and governance requirements. Evaluate emerging technologies and make recommendations for platform evolution. Participate in architecture design reviews and provide technical leadership on complex initiatives. Complementary AI Based Skills: LLM Orchestration & Agentic Workflows: Experience in designing, building, and deploying Large Language Model (LLM) orchestration frameworks (e.g., LangChain, Temporal, or custom agentic loops) to coordinate multi-step diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root-cause analysis, anomaly detection, and semantic log clustering. Self-Healing Infrastructure Engineering: Experience designing closed-loop, self-healing systems that autonomously execute recovery actions (e.g., traffic shifting, automated rollbacks, or service restarts) with built-in verification and safety guardrails. AI-Driven FinOps & Resource Optimization: Deep understanding of applying machine learning and predictive analytics to dynamically right-size cloud resources, manage spot instances, and optimize data platform workloads. Predictive Capacity Planning: Ability to design algorithms that forecast workload demands and proactively scale infrastructure to prevent over-provisioning while maintaining strict SLAs. Complementary Behaviours: Toil-Reduction Mindset: A relentless focus on eliminating repetitive operational support and engineering friction by shifting platform operations from reactive troubleshooting to autonomous mitigation. Risk-Aware Automation: Demonstrates a disciplined approach to safety by implementing strict confidence thresholds, validation loops, and human-in-the-loop fallbacks for autonomous AI actions. Value-Oriented Engineering: Treats cost optimization as a first-class architectural metric, aligning infrastructure spend directly with business value and platform efficiency. Impact-Preserving Innovation: Executes large-scale optimization initiatives with a meticulous, risk-mitigated approach, ensuring zero disruption to production environments or developer velocity. Basic Qualifications: 10+ years of experience in software engineering or infrastructure engineering. Deep hands-on expertise with AWS services (EC2, EKS, Lambda, S3, IAM, VPC, CloudFormation, CDK, etc.). Strong background in platform engineering, building internal developer platforms, or infrastructure tooling. Experience designing and operating large-scale distributed systems. Proficiency with infrastructure-as-code tools (Terraform, CloudFormation). Strong understanding of containerization and orchestration (Docker, Kubernetes). Experience with CI/CD pipelines and DevOps practices. Knowledge of networking, security, and identity management in cloud environments. Excellent communication skills with the ability to influence technical decisions across teams. Experience working in regulated industries is a plus. Preferred Qualifications: AWS certifications (Solutions Architect Professional, DevOps Engineer, etc.). Experience building self-service platforms for development teams. Familiarity with observability and monitoring tools (Prometheus, Grafana, Datadog, CloudWatch). Background in financial services or other highly regulated environments. About Goldman Sachs At Goldman Sachs, we commit our people, capital and ideas to help our clients, shareholders and the communities we serve to grow. Founded in 1869, we are a leading global investment banking, securities and investment management firm. Headquartered in New York, we maintain offices around the world. We believe who you are makes you better at what you do. We're committed to fostering and advancing diversity and inclusion in our own workplace and beyond by ensuring every individual within our firm has several opportunities to grow professionally and personally, from our training and development opportunities and firmwide networks to benefits, wellness and personal finance offerings and mindfulness programs. Learn more about our culture, benefits, and people at . We're committed to finding reasonable accommodation for candidates with special needs or disabilities during our recruiting process. Learn more:
What We Do At the company, our Engineers don't just make things - we make things possible. Change the world by connecting people and capital with ideas. Solve the most challenging and pressing engineering problems for our clients. Join our engineering teams that build massively scalable software and systems, architect low latency infrastructure solutions, proactively guard against cyber threats, and leverage machine learning alongside financial engineering to continuously turn data into action. Create new businesses, transform finance, and explore a world of opportunity at the speed of markets. the company Engineers are innovators and problem-solvers, building solutions in Artificial Intelligence, risk management, big data, mobile and more. Cloud Engineering & Architecture (CE&A) As part of Core Engineering at the company, the CE&A team is responsible for enabling the use of public cloud services across the firm. You will be working as part of a multi-disciplinary team responsible for researching, architecting and building a cutting edge platform that enable the company Engineering teams to deploy and manage services in public cloud safely and securely. The organization is seeking highly collaborative, creative, and intellectually curious engineers who are passionate about developing and implementing cutting edge cloud computing and AI solutions. The ideal candidate will thrive in a DevOps culture and contribute to customer centric product development. They will work closely with cross functional teams, and will be creative collaborators who evolve, adapt to change and thrive in a fast paced global environment. Responsibilities And Qualifications: We are looking for a senior technical leader to join our Cloud Engineering & Architecture team and play a pivotal role in enabling the firm to maximize its use of cloud infrastructure. This is a hands on leadership position requiring deep technical expertise, strategic thinking, and the ability to drive large scale platform initiatives from conception to delivery. Key Responsibilities: Design, develop, and operationalize enterprise grade cloud platform capabilities. Architect scalable, resilient, and secure infrastructure solutions on AWS. Architect and operationalize autonomous AI based, self healing infrastructure Define technical standards, best practices, and reference architectures for cloud adoption across the firm. Partner with engineering teams to enable seamless migration and modernization of workloads to the cloud. Drive automation and infrastructure as code practices to improve operational efficiency. Drive AI powered FinOps and predictive resource optimization Mentor and guide engineers across teams, raising the overall technical bar. Collaborate with security, networking, and compliance teams to ensure platform meets regulatory and governance requirements. Evaluate emerging technologies and make recommendations for platform evolution. Participate in architecture design reviews and provide technical leadership on complex initiatives. Complementary AI Based Skills: LLM Orchestration & Agentic Workflows: Experience in designing, building, and deploying Large Language Model (LLM) orchestration frameworks (e.g., LangChain, Temporal, or custom agentic loops) to coordinate multi step diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root cause analysis, anomaly detection, and semantic log clustering. Self Healing Infrastructure Engineering: Experience designing closed loop, self healing systems that autonomously execute recovery actions (e.g., traffic shifting, automated rollbacks, or service restarts) with built in verification and safety guardrails. AI Driven FinOps & Resource Optimization: Deep understanding of applying machine learning and predictive analytics to dynamically right size cloud resources, manage spot instances, and optimize data platform workloads. Predictive Capacity Planning: Ability to design algorithms that forecast workload demands and proactively scale infrastructure to prevent over provisioning while maintaining strict SLAs. Complementary Behaviours: Toil Reduction Mindset: A relentless focus on eliminating repetitive operational support and engineering friction by shifting platform operations from reactive troubleshooting to autonomous mitigation. Risk Aware Automation: Demonstrates a disciplined approach to safety by implementing strict confidence thresholds, validation loops, and human in the loop fallbacks for autonomous AI actions. Value Oriented Engineering: Treats cost optimization as a first class architectural metric, aligning infrastructure spend directly with business value and platform efficiency. Impact Preserving Innovation: Executes large scale optimization initiatives with a meticulous, risk mitigated approach, ensuring zero disruption to production environments or developer velocity. Basic Qualifications: 10+ years of experience in software engineering or infrastructure engineering. Deep hands on expertise with AWS services (EC2, EKS, Lambda, S3, IAM, VPC, CloudFormation, CDK, etc.). Strong background in platform engineering, building internal developer platforms, or infrastructure tooling. Experience designing and operating large scale distributed systems. Proficiency with infrastructure as code tools (Terraform, CloudFormation). Strong understanding of containerization and orchestration (Docker, Kubernetes). Experience with CI/CD pipelines and DevOps practices. Knowledge of networking, security, and identity management in cloud environments. Excellent communication skills with the ability to influence technical decisions across teams. Experience working in regulated industries is a plus. Preferred Qualifications: AWS certifications (Solutions Architect Professional, DevOps Engineer, etc.). Experience building self service platforms for development teams. Familiarity with observability and monitoring tools (Prometheus, Grafana, Datadog, CloudWatch). Background in financial services or other highly regulated environments. About the company At the company, we commit our people, capital and ideas to help our clients, shareholders and the communities we serve to grow. Founded in 1869, we are a leading global investment banking, securities and investment management firm. Headquartered in New York, we maintain offices around the world. We believe who you are makes you better at what you do. We're committed to fostering and advancing diversity and inclusion in our own workplace and beyond by ensuring every individual within our firm has several opportunities to grow professionally and personally, from our training and development opportunities and firmwide networks to benefits, wellness and personal finance offerings and mindfulness programs. about our culture, benefits, and people at We're committed to finding reasonable accommodation for candidates with special needs or disabilities during our recruiting process. : We Offer Best-In-Class Benefits Healthcare & Medical Insurance We offer a wide range of health and welfare programs that vary depending on office location. These generally include medical, dental, short term disability, long term disability, life, accidental death, labor accident and business travel accident insurance. Holiday & Vacation Policies We offer competitive vacation policies based on employee level and office location. We promote time off from work to recharge by providing generous vacation entitlements and a minimum of three weeks expected vacation usage each year. Financial Wellness & Retirement We assist employees in saving and planning for retirement, offer financial support for higher education, and provide a number of benefits to help employees prepare for the unexpected. We offer live financial education and content on a variety of topics to address the spectrum of employees' priorities. Health Services We offer a medical advocacy service for employees and family members facing critical health situations, and counseling and referral services through the Employee Assistance Program (EAP). We provide Global Medical, Security and Travel Assistance and a Workplace Ergonomics Program. We also offer state of the art on site health centers in certain offices. Fitness To encourage employees to live a healthy and active lifestyle, some of our offices feature on site fitness centers. For eligible employees we typically reimburse fees paid for a fitness club membership or activity (up to a pre approved amount). Child Care & Family Care We offer on site child care centers that provide full time and emergency back up care, as well as mother and baby rooms and homework rooms. In every office, we provide advice and counseling services, expectant parent resources and transitional programs for parents returning from parental leave. Adoption, surrogacy, egg donation and egg retrieval stipends are also available. Benefits at the company Read more about the full suite of class leading benefits our firm has to offer.
26/07/2026
Full time
What We Do At the company, our Engineers don't just make things - we make things possible. Change the world by connecting people and capital with ideas. Solve the most challenging and pressing engineering problems for our clients. Join our engineering teams that build massively scalable software and systems, architect low latency infrastructure solutions, proactively guard against cyber threats, and leverage machine learning alongside financial engineering to continuously turn data into action. Create new businesses, transform finance, and explore a world of opportunity at the speed of markets. the company Engineers are innovators and problem-solvers, building solutions in Artificial Intelligence, risk management, big data, mobile and more. Cloud Engineering & Architecture (CE&A) As part of Core Engineering at the company, the CE&A team is responsible for enabling the use of public cloud services across the firm. You will be working as part of a multi-disciplinary team responsible for researching, architecting and building a cutting edge platform that enable the company Engineering teams to deploy and manage services in public cloud safely and securely. The organization is seeking highly collaborative, creative, and intellectually curious engineers who are passionate about developing and implementing cutting edge cloud computing and AI solutions. The ideal candidate will thrive in a DevOps culture and contribute to customer centric product development. They will work closely with cross functional teams, and will be creative collaborators who evolve, adapt to change and thrive in a fast paced global environment. Responsibilities And Qualifications: We are looking for a senior technical leader to join our Cloud Engineering & Architecture team and play a pivotal role in enabling the firm to maximize its use of cloud infrastructure. This is a hands on leadership position requiring deep technical expertise, strategic thinking, and the ability to drive large scale platform initiatives from conception to delivery. Key Responsibilities: Design, develop, and operationalize enterprise grade cloud platform capabilities. Architect scalable, resilient, and secure infrastructure solutions on AWS. Architect and operationalize autonomous AI based, self healing infrastructure Define technical standards, best practices, and reference architectures for cloud adoption across the firm. Partner with engineering teams to enable seamless migration and modernization of workloads to the cloud. Drive automation and infrastructure as code practices to improve operational efficiency. Drive AI powered FinOps and predictive resource optimization Mentor and guide engineers across teams, raising the overall technical bar. Collaborate with security, networking, and compliance teams to ensure platform meets regulatory and governance requirements. Evaluate emerging technologies and make recommendations for platform evolution. Participate in architecture design reviews and provide technical leadership on complex initiatives. Complementary AI Based Skills: LLM Orchestration & Agentic Workflows: Experience in designing, building, and deploying Large Language Model (LLM) orchestration frameworks (e.g., LangChain, Temporal, or custom agentic loops) to coordinate multi step diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root cause analysis, anomaly detection, and semantic log clustering. Self Healing Infrastructure Engineering: Experience designing closed loop, self healing systems that autonomously execute recovery actions (e.g., traffic shifting, automated rollbacks, or service restarts) with built in verification and safety guardrails. AI Driven FinOps & Resource Optimization: Deep understanding of applying machine learning and predictive analytics to dynamically right size cloud resources, manage spot instances, and optimize data platform workloads. Predictive Capacity Planning: Ability to design algorithms that forecast workload demands and proactively scale infrastructure to prevent over provisioning while maintaining strict SLAs. Complementary Behaviours: Toil Reduction Mindset: A relentless focus on eliminating repetitive operational support and engineering friction by shifting platform operations from reactive troubleshooting to autonomous mitigation. Risk Aware Automation: Demonstrates a disciplined approach to safety by implementing strict confidence thresholds, validation loops, and human in the loop fallbacks for autonomous AI actions. Value Oriented Engineering: Treats cost optimization as a first class architectural metric, aligning infrastructure spend directly with business value and platform efficiency. Impact Preserving Innovation: Executes large scale optimization initiatives with a meticulous, risk mitigated approach, ensuring zero disruption to production environments or developer velocity. Basic Qualifications: 10+ years of experience in software engineering or infrastructure engineering. Deep hands on expertise with AWS services (EC2, EKS, Lambda, S3, IAM, VPC, CloudFormation, CDK, etc.). Strong background in platform engineering, building internal developer platforms, or infrastructure tooling. Experience designing and operating large scale distributed systems. Proficiency with infrastructure as code tools (Terraform, CloudFormation). Strong understanding of containerization and orchestration (Docker, Kubernetes). Experience with CI/CD pipelines and DevOps practices. Knowledge of networking, security, and identity management in cloud environments. Excellent communication skills with the ability to influence technical decisions across teams. Experience working in regulated industries is a plus. Preferred Qualifications: AWS certifications (Solutions Architect Professional, DevOps Engineer, etc.). Experience building self service platforms for development teams. Familiarity with observability and monitoring tools (Prometheus, Grafana, Datadog, CloudWatch). Background in financial services or other highly regulated environments. About the company At the company, we commit our people, capital and ideas to help our clients, shareholders and the communities we serve to grow. Founded in 1869, we are a leading global investment banking, securities and investment management firm. Headquartered in New York, we maintain offices around the world. We believe who you are makes you better at what you do. We're committed to fostering and advancing diversity and inclusion in our own workplace and beyond by ensuring every individual within our firm has several opportunities to grow professionally and personally, from our training and development opportunities and firmwide networks to benefits, wellness and personal finance offerings and mindfulness programs. about our culture, benefits, and people at We're committed to finding reasonable accommodation for candidates with special needs or disabilities during our recruiting process. : We Offer Best-In-Class Benefits Healthcare & Medical Insurance We offer a wide range of health and welfare programs that vary depending on office location. These generally include medical, dental, short term disability, long term disability, life, accidental death, labor accident and business travel accident insurance. Holiday & Vacation Policies We offer competitive vacation policies based on employee level and office location. We promote time off from work to recharge by providing generous vacation entitlements and a minimum of three weeks expected vacation usage each year. Financial Wellness & Retirement We assist employees in saving and planning for retirement, offer financial support for higher education, and provide a number of benefits to help employees prepare for the unexpected. We offer live financial education and content on a variety of topics to address the spectrum of employees' priorities. Health Services We offer a medical advocacy service for employees and family members facing critical health situations, and counseling and referral services through the Employee Assistance Program (EAP). We provide Global Medical, Security and Travel Assistance and a Workplace Ergonomics Program. We also offer state of the art on site health centers in certain offices. Fitness To encourage employees to live a healthy and active lifestyle, some of our offices feature on site fitness centers. For eligible employees we typically reimburse fees paid for a fitness club membership or activity (up to a pre approved amount). Child Care & Family Care We offer on site child care centers that provide full time and emergency back up care, as well as mother and baby rooms and homework rooms. In every office, we provide advice and counseling services, expectant parent resources and transitional programs for parents returning from parental leave. Adoption, surrogacy, egg donation and egg retrieval stipends are also available. Benefits at the company Read more about the full suite of class leading benefits our firm has to offer.
At Moneybox, our mission is to give everyone the means to get more out of life. We're guided by our belief that wealth isn't about the money, it's about the means to more - more freedom, opportunities, possibilities, and peace of mind. Moneybox is an award winning wealth management platform, helping over one and a half million people build wealth throughout their lives, whether they're saving and investing, buying their first home, or planning for retirement. This is a hybrid role. Our office is in London, by the Oxo Tower. Job brief This role is in our Tech Ops Engineering team that operates our cloud hosted services. You will be working with people throughout Moneybox to design and deliver platform capabilities, drive operational excellence, support the live service, and shape how we do things. We're looking for someone who thrives across solution architecture, hands on engineering, and owning their solutions from concept through to production - and who can bring others along with them. At senior level, we expect you to lead technical direction within the team, make sound decisions under pressure, and actively improve the systems and practices around you. We're actively embracing AI, and are keen to hear from those with experience or enthusiasm for this new approach. This team offers and runs the following services: System Reliability: Maintains and enhances the scalability, availability, security and resilience of systems hosted in the Azure cloud. Collaborative Development: Provides infrastructure patterns and practices guidance to squads as they develop new services, and continuously improve the IaC, policies and guardrails that allow them to operate independently. Advanced Technical Support: Works alongside our engineering teams to handle complex technical inquiries related to our live services, ensuring service issues are resolved swiftly, with minimal impact. Our Tech Stack Azure: App Services, Functions, Service Bus, Event Hub, CosmosDB, Redis, SQL Server, Databricks, Key Vault, AKS (Kubernetes), Networking, Managed Identity Infrastructure as Code (Terraform / Terraform Cloud) CloudFlare GitHub, GitHub Actions, Azure DevOps Pipelines Datadog C#, .NET Core / .NET Framework (being phased out) REST APIs, Hangfire, MediatR, Entity Framework, Mass Transit, xUnit/NUnit Node, Next.js, React PowerShell, Bash What you'll do Lead the design, build, and operation of cloud infrastructure, platform services, and ops tooling that underpin our live service and empower engineering teams. Own and evolve monitoring and alerting across production systems - ensuring dashboards and alerts are high signal and actionable. Lead incident response during production issues. Coordinate across teams, drive resolution, and ensure thorough post incident reviews with meaningful follow up actions. Investigate and resolve complex technical issues across application, infrastructure, data, and networking layers - performing root cause analysis and implementing durable fixes, not just patches. Maintain, evolve, and champion our Infrastructure as Code (Terraform) to ensure environments are reproducible, auditable, and version controlled. Improve CI/CD pipelines and deployment practices to increase engineering velocity while maintaining deployment safety and rollback capability. Proactively manage cloud capacity and cost - monitoring spend, identifying optimisation opportunities, and contributing to FinOps practices including tagging, budgeting, and cost allocation. Ensure infrastructure and platform services meet security, regulatory, and compliance requirements. Implement and maintain controls around network segmentation, access management, secrets handling, and vulnerability patching. Support audit and governance processes with evidence and documentation. Drive automation of operational toil and repetitive tasks - if you're doing it more than twice, automate it. Contribute to production readiness standards - every change should be observable, reversible, and well tested before it reaches customers. Foster a knowledge sharing environment with thorough documentation, runbooks, and a teamwork oriented culture. Support the wider business to meet their goals where major infrastructure or service change is required. Support, coach, and mentor other team members. Raise the technical bar through pairing, code review, and leading by example. Stay abreast of and (where necessary) apply the latest emerging technologies relevant to systems engineering and cloud infrastructure. Who you are Passionate about platform reliability, operational excellence, and building shared ownership of production systems across the wider engineering team. Comfortable leading incident response under pressure and making sound technical decisions with incomplete information. Excited about being part of a fast growing company that's trying to make a positive mark on the world. A driven, ambitious self starter who takes ownership and sees things through. Collaborative attitude - you enjoy working individually as well as within a team, and you naturally bring people along with your ideas. Pragmatic - you favour simplicity, avoid over engineering, and know when "good enough" is the right call. Can embrace our ALOT values. Knows how to have fun whilst maintaining a professional outlook. Essential Skills A degree in Computer Science or relevant experience Proven track record in a similar role Strong understanding of Cloud Infrastructure (even better if it's Microsoft Azure) Infrastructure as Code (Terraform) Web Application Security (e.g. CloudFlare) Web and API scalability and performance Distributed systems Build and Release Pipelines (e.g. Azure DevOps, Github Actions) Strong analytical and problem solving skills Comfortable working within a live Cloud environment
26/07/2026
Full time
At Moneybox, our mission is to give everyone the means to get more out of life. We're guided by our belief that wealth isn't about the money, it's about the means to more - more freedom, opportunities, possibilities, and peace of mind. Moneybox is an award winning wealth management platform, helping over one and a half million people build wealth throughout their lives, whether they're saving and investing, buying their first home, or planning for retirement. This is a hybrid role. Our office is in London, by the Oxo Tower. Job brief This role is in our Tech Ops Engineering team that operates our cloud hosted services. You will be working with people throughout Moneybox to design and deliver platform capabilities, drive operational excellence, support the live service, and shape how we do things. We're looking for someone who thrives across solution architecture, hands on engineering, and owning their solutions from concept through to production - and who can bring others along with them. At senior level, we expect you to lead technical direction within the team, make sound decisions under pressure, and actively improve the systems and practices around you. We're actively embracing AI, and are keen to hear from those with experience or enthusiasm for this new approach. This team offers and runs the following services: System Reliability: Maintains and enhances the scalability, availability, security and resilience of systems hosted in the Azure cloud. Collaborative Development: Provides infrastructure patterns and practices guidance to squads as they develop new services, and continuously improve the IaC, policies and guardrails that allow them to operate independently. Advanced Technical Support: Works alongside our engineering teams to handle complex technical inquiries related to our live services, ensuring service issues are resolved swiftly, with minimal impact. Our Tech Stack Azure: App Services, Functions, Service Bus, Event Hub, CosmosDB, Redis, SQL Server, Databricks, Key Vault, AKS (Kubernetes), Networking, Managed Identity Infrastructure as Code (Terraform / Terraform Cloud) CloudFlare GitHub, GitHub Actions, Azure DevOps Pipelines Datadog C#, .NET Core / .NET Framework (being phased out) REST APIs, Hangfire, MediatR, Entity Framework, Mass Transit, xUnit/NUnit Node, Next.js, React PowerShell, Bash What you'll do Lead the design, build, and operation of cloud infrastructure, platform services, and ops tooling that underpin our live service and empower engineering teams. Own and evolve monitoring and alerting across production systems - ensuring dashboards and alerts are high signal and actionable. Lead incident response during production issues. Coordinate across teams, drive resolution, and ensure thorough post incident reviews with meaningful follow up actions. Investigate and resolve complex technical issues across application, infrastructure, data, and networking layers - performing root cause analysis and implementing durable fixes, not just patches. Maintain, evolve, and champion our Infrastructure as Code (Terraform) to ensure environments are reproducible, auditable, and version controlled. Improve CI/CD pipelines and deployment practices to increase engineering velocity while maintaining deployment safety and rollback capability. Proactively manage cloud capacity and cost - monitoring spend, identifying optimisation opportunities, and contributing to FinOps practices including tagging, budgeting, and cost allocation. Ensure infrastructure and platform services meet security, regulatory, and compliance requirements. Implement and maintain controls around network segmentation, access management, secrets handling, and vulnerability patching. Support audit and governance processes with evidence and documentation. Drive automation of operational toil and repetitive tasks - if you're doing it more than twice, automate it. Contribute to production readiness standards - every change should be observable, reversible, and well tested before it reaches customers. Foster a knowledge sharing environment with thorough documentation, runbooks, and a teamwork oriented culture. Support the wider business to meet their goals where major infrastructure or service change is required. Support, coach, and mentor other team members. Raise the technical bar through pairing, code review, and leading by example. Stay abreast of and (where necessary) apply the latest emerging technologies relevant to systems engineering and cloud infrastructure. Who you are Passionate about platform reliability, operational excellence, and building shared ownership of production systems across the wider engineering team. Comfortable leading incident response under pressure and making sound technical decisions with incomplete information. Excited about being part of a fast growing company that's trying to make a positive mark on the world. A driven, ambitious self starter who takes ownership and sees things through. Collaborative attitude - you enjoy working individually as well as within a team, and you naturally bring people along with your ideas. Pragmatic - you favour simplicity, avoid over engineering, and know when "good enough" is the right call. Can embrace our ALOT values. Knows how to have fun whilst maintaining a professional outlook. Essential Skills A degree in Computer Science or relevant experience Proven track record in a similar role Strong understanding of Cloud Infrastructure (even better if it's Microsoft Azure) Infrastructure as Code (Terraform) Web Application Security (e.g. CloudFlare) Web and API scalability and performance Distributed systems Build and Release Pipelines (e.g. Azure DevOps, Github Actions) Strong analytical and problem solving skills Comfortable working within a live Cloud environment