Overview
Lead the Site Reliability Engineering (SRE) practice, driving the transformation from reactive operations to proactive, engineering led reliability. Own the definition and enforcement of non functional requirements (NFRs) using FMEA based resiliency frameworks, and champion observability, self healing automation, automated incident management, and database operations automation.
Key Responsibilities
- Define and enforce NFRs for performance, scalability, availability, fault tolerance, and cost efficiency using FMEA based failure analysis.
- Design and implement self healing automation for known failure patterns, reducing human intervention and on call burden by 50%+.
- Build comprehensive observability stacks (metrics, logs, traces) with ML driven anomaly detection and AIOps capabilities.
- Lead automated incident management: detection, triage, escalation, remediation, and post incident review automation.
- Drive database automation: automated provisioning, release management (UK focus), backup/restore, and operational request workflows for all database operations.
- Define and track SLIs, SLOs, and error budgets across all critical services, using them to balance reliability with feature velocity.
- Conduct chaos engineering exercises and game days to validate resiliency and uncover hidden failure modes.
- Mentor two SRE engineers, establish engineering standards, and build a culture of reliability and continuous improvement.
- Collaborate with Platform Engineering and Cloud teams to embed reliability into infrastructure and deployment pipelines.
Technical Skills & Expertise
- Expert level observability: Prometheus, Grafana, ELK/OpenSearch, Jaeger/Zipkin, Datadog, or Dynatrace.
- Strong experience with AIOps and ML driven monitoring: PagerDuty, Moogsoft, BigPanda, or custom ML pipelines.
- Deep knowledge of FMEA, fault tree analysis, and chaos engineering tools (Gremlin, LitmusChaos, Chaos Monkey).
- Database automation: strong SQL skills plus experience with automated DB provisioning, migration tools (Liquibase, Flyway), and DB release pipelines.
- Proficiency in automation and scripting: Python, Go, Bash, with experience building self healing runbooks.
- Infrastructure knowledge: Kubernetes, cloud platforms (AWS/Azure/GCP), networking, and storage systems.
- CI/CD and release engineering: Jenkins, GitLab CI, Spinnaker, ArgoCD for integrated DB and application releases.
- Cost management: experience with FinOps principles, resource optimisation, and cloud spend analysis.
Soft Skills & Competencies
- Strong leadership and mentoring ability-coaches and develops junior engineers.
- Excellent stakeholder management and communication skills across all levels.
- Strategic thinker who balances technical depth with business outcomes.
- Proven ability to drive change, influence without authority, and build consensus.
- Strong analytical and problem solving mindset with attention to detail.
- Ability to manage competing priorities across multiple workstreams simultaneously.
Qualifications & Experience
- 7+ years in SRE, DevOps, or production engineering with 3+ years in a senior or lead capacity.
- Proven track record of improving availability, reducing MTTR, and implementing self healing at scale.
- Experience managing or automating database operations in enterprise environments.
- Relevant certifications preferred: CKA, AWS DevOps Professional, Azure DevOps Expert, SRE Foundation.
- Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience).
Desirable / Nice to Have
- Published work or conference talks on SRE, observability, or chaos engineering.
- Experience with service mesh (Istio) and distributed tracing at scale.
- Background in financial services or regulated industry.
EEO Statement
We're an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.