Observability Architect

  • Jobtailor
  • Bristol, Gloucestershire
  • 22/07/2026
Full time Information Technology Telecommunications Testing

Job Description

Job Responsibilities
  • Lead the assessment, design, and optimisation of the observability strategy for the co-location migration programme.
  • Review the current observability architecture across infrastructure, networks, middleware, databases, and applications.
  • Assess existing logging, metrics, distributed tracing, and monitoring capabilities to determine readiness for the co-location migration.
  • Develop an observability strategy that supports both migration activities and long-term operational support.
  • Recommend enhancements or platform uplifts where current tooling does not provide sufficient visibility or resilience.
  • Analyse telemetry, monitoring data, dashboards, and operational trends from completed migration waves.
  • Establish performance baselines for compute, storage, networking, application response times, and transaction throughput.
  • Identify recurring operational issues and use historical insights to improve migration readiness.
  • Define measurable service health indicators to compare pre- and post-migration performance.
  • Design comprehensive monitoring for the tightly-coupled monolithic application estate, with particular emphasis on latency-sensitive interdependencies.
  • Create real-time dashboards that provide operational visibility across infrastructure, middleware, databases, messaging, and application components.
  • Ensure end-to-end transaction tracing is available to rapidly identify bottlenecks and service degradation.
  • Validate monitoring coverage prior to each migration wave.
  • Review and standardise centralised logging across all migrated environments.
  • Ensure consistent log formats, metadata, correlation IDs, and traceability across systems.
  • Validate log ingestion, retention policies, indexing, and search performance.
  • Ensure operational teams can rapidly investigate incidents using correlated logs and distributed traces.
  • Review and optimise alert thresholds to minimise both missed events and unnecessary alert noise.
  • Implement intelligent alerting aligned to business services and critical customer journeys.
  • Define migration-specific alerting for infrastructure failures, application degradation, latency increases, replication issues, and capacity constraints.
  • Support operational readiness activities including rehearsals and production cutover monitoring.
  • Ensure observability solutions meet financial services regulatory requirements for auditability, log retention, security, and data governance.
  • Validate access controls and security monitoring for observability platforms.
  • Support evidence gathering for internal governance, audit, and regulatory reviews.
  • Evaluate the suitability of existing observability platforms and recommend improvements where required.
  • Assess opportunities to improve automation, anomaly detection, service health monitoring, and predictive alerting.
  • Define standards and best practices for observability across future migration phases.
  • Work closely with Infrastructure Architects, Application Architects, Platform Engineering, Security, Operations, and Migration teams.
  • Provide technical guidance during migration planning, testing, dress rehearsals, and production cutovers.
  • Produce architecture documentation, monitoring standards, operational runbooks, and knowledge transfer materials.
Requirements
  • Extensive experience designing enterprise observability solutions within large-scale infrastructure or data centre migration programmes.
  • Strong knowledge of metrics, logging, distributed tracing, and application performance monitoring (APM).
  • Experience monitoring latency-sensitive, business-critical enterprise applications.
  • Strong understanding of infrastructure, virtualisation, networking, storage, databases, and middleware monitoring.
  • Experience implementing centralised logging and observability best practices.
  • Knowledge of financial services operational resilience, audit, and regulatory requirements.
  • Ability to analyse complex operational telemetry and identify performance bottlenecks.
  • Excellent stakeholder management and communication skills.