> Markdown version of [/jobs/ext/2447288-senior-platform-sre](https://www.wearedevelopers.com/jobs/ext/2447288-senior-platform-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Platform SRE - **Company:** The IG Group - **Location:** London, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Amazon Web Services, Build Automation, Common Lisp Object Systems, Cloud Computing, Code Review, Continuous Integration, Distributed Systems, Factor Analysis, Fault Tolerance, Python (Programming Language), Reliability Engineering, Software Engineering, Datadog, Grafana, Reliability of Systems, Kubernetes, Deployment Automation, Hashicorp, Virtual Agents, Terraform, Software Version Control, Dynatrace, Pagerduty, Servicenow - **Published:** August 4, 2026 - **Apply:** https://uk.indeed.com/viewjob?jk=43dc19a6d5e1ec99 ## About the Role Extensive working experience in all the below mentioned areas. Essential Technical Skills * Observability and instrumentation: hands-on OpenTelemetry experience (spans, metrics, traces, context propagation) and production use of Honeycomb, Datadog, Dynatrace, or Grafana; able to instrument Java or Python services directly . * SLOs and error budgets: proven track record designing customer-meaningful SLIs, setting error budgets, configuring multi-window burn-rate alerts, and working with development teams on reliability measurement * CI/CD and release engineering: experience building pipelines with safety mechanisms: blue/green and canary releases, automated rollback, and DORA metrics integration * Container orchestration: Kubernetes (EKS, AKS, or GKE) required ; HashiCorp Nomad is a strong advantage on IG's hybrid estate; solid understanding of cloud networking and IaC (Terraform preferred) * Software engineering: production-quality coding in Java and/or Python; comfortable contributing to application codebases to implement reliability patterns, not just configuring infrastructure around them * Distributed systems: strong understanding of how large-scale systems fail and how to make them fail safely; circuit breakers, bulkheads, idempotency, graceful degradation, and load-shedding; high-throughput, low-latency environments preferred * Incident management: on-call experience on production systems, blameless PIR facilitation, contributing-factor analysis, and driving action items to closure; PagerDuty and ServiceNow familiarity helpful * Chaos engineering: experience designing and executing hypothesis-driven experiments with blast-radius controls and gap-to-impact-tolerance analysis; AWS FIS, Gremlin, or equivalent * Community and standards: at ease in a guild or community-of-practice model; comfortable writing RFCs, presenting at engineering forums, and building standards that others will adopt Experience Requirements * Track record in high-throughput, production environments (financial services, trading platforms, or similar mission-critical systems preferred) * Demonstrated ability to improve system reliability and performance at scale * Experience working collaboratively with development teams to implement observability and reliability improvements * Strong troubleshooting skills in distributed systems environments Core Competencies * Systems thinking approach to problem-solving * Excellent communication skills for cross-functional collaboration and technical enablement * Ability to balance hands-on development work with operational responsibilities * Strong bias toward automation and eliminating manual toil ## Description You will be a hands-on technical contributor at the heart of the Platform SRE team, owning pieces of the reliability platform that hundreds of engineers depend on. You will work at the intersection of software engineering, observability, and systems reliability, turning reliability from a reactive concern into a proactive engineering discipline., * Implement comprehensive monitoring and observability using OpenTelemetry and distributed tracing . Maintain SLO , error budgets and burn-rate tracking * Establish and maintain 24/7 operational readiness including automated deployments, blue/green releases, and zero-downtime patching strategies * Engineer self-healing capabilities: auto-remediation, error-budget-gated rollback, and automated traffic rerouting * Design and run chaos experiments across the AWS estate, turning severe-but-plausible failure scenarios into engineering improvements * Build automation tools and CI/CD pipelines that embed reliability practices , while a pply ing software engineering discipline including version control, code reviews, and testing . * Contribute to the SRE AI agent, IG's agentic tooling for incident investigation and reliability review, built on AWS frontier models * Mentor junior SREs and Reliability Champions on reliability patterns and production engineering discipline Set and uphold standards * Author and evolve the SRE standards that underpin the Guild: SLO methodology , error budget policy, observability instrumentation guide, and Production Readiness Review (PRR) checklist * Mentor developers on reliability patterns including circuit breakers, retry logic, and fault tolerance * Work with development teams and Reliability Champions to design SLOs on customer journeys rather than per-service . * Assist and guide teams in system design, capacity planning, architectural reviews and clos ing observability gaps . Own incident response and learning * Facilitate blameless post-incident reviews (PIRs) within five working days using contributing-factor methodology * Maintain the Lessons Register, track remediation actions to closure, and surface patterns across incidents quarterly ## Related Videos - [AI in Leadership: How Technology is Reshaping Executive Roles](https://www.wearedevelopers.com/videos/1705-ai-in-leadership-how-technology-is-reshaping-executive-roles) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Best Companies to work for in London: Top 25 Companies in 2023](https://www.wearedevelopers.com/magazine/187-best-companies-to-work-for-in-london-top-25-companies-in-2023) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Software Engineer Salary London](https://www.wearedevelopers.com/magazine/252-software-engineer-salary-london) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)