> Markdown version of [/jobs/ext/3332267-lead-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3332267-lead-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Site Reliability Engineer - **Company:** JPMorgan Chase & Co. - **Location:** London, UK - **Experience:** Expert - **Salary:** £62,000.0 - £102,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Algorithmic Trading, Automation of Tests, Software Bug Management, Cloud Computing, Cloud Engineering, Cyber Security, Continuous Integration, Disaster Recovery, Distributed Systems, Python (Programming Language), Operational Data Store, Oracle (Applications), Performance Tuning, Systems Development Life Cycle, Reliability Engineering, Software Engineering, Grafana, Kotlin, Event Driven Architecture, Low Latency, Influxdb, Deployment Automation, Production Code, Apache Kafka, Splunk, Dynatrace, Microservices, Oracledb - **Published:** September 9, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5875482339 ## About the Role * Strong hands-on experience in front-office trading environments or similarly high-pressure, low-latency domains. * Proficiency with SRE tooling and techniques, including FIX messaging, Kafka, Grafana, Splunk, ITRS Geneos, Dynatrace, InfluxDB, MQ, and Oracle DB. * Demonstrated experience using enterprise-authorized AI capabilities to improve SRE workflows, with strong validation habits and awareness of data sensitivity. * Ability to evaluate AI-assisted operational recommendations for correctness and risk, define guardrails for team usage, and ensure outcomes align with resiliency and security expectations. * Deep knowledge of reliability engineering principles, including SLIs/SLOs, real-time telemetry, disaster recovery planning, capacity planning, and performance tuning. * Experience designing and implementing observability frameworks for mission-critical systems. * Proven ability to lead incident response and drive long-term remediation. * Solid programming skills in Python, Java, or Kotlin, with the ability to contribute production-grade code. * Experience with microservices, distributed systems, and event-driven architectures. * Strong understanding of CI/CD pipelines, automated testing, and deployment strategies. * Comfort interacting directly with traders and senior stakeholders. * Excellent communication skills, especially in translating technical issues into business impact. * Ability to operate calmly and decisively in high-pressure situations. * Strong leadership presence with a collaborative mindset. ## Description * Engage daily with traders across asset classes to understand workflows, pain points, and reliability priorities. * Act as a trusted engineering partner to the desk, ensuring systems are stable, performant, and aligned with business needs. * Support live trading environments, including incident response, root cause analysis, and post-mortem leadership. * Participate as a core member of the software engineering team in daily standups and design discussions. * Contribute directly to the codebase to implement reliability improvements, performance optimisations, bug fixes, and automation. * Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self-healing workflows, and resilience engineering. * Use enterprise-authorized AI capabilities to accelerate major-incident triage, troubleshooting, and post-incident analysis, while validating outputs and handling operational data securely. * Drive reuse-first adoption of AI-assisted reliability workflows across SDLC and toolchain practices, ensuring traceability, auditability, resiliency, and security controls. * Drive improvements in latency, throughput, and stability across high-volume trading applications. * Build and maintain tooling for monitoring, alerting, and distributed tracing across global environments. * Operate within a globally distributed engineering and trading organization, collaborating with teams in EMEA, US, and APAC. * Partner with infrastructure, networking, cloud engineering, and cybersecurity teams to ensure end-to-end reliability. Technologies: * AI * CI/CD * Cloud * Dynatrace * Embedded * Grafana * Support * Java * Kafka * Kotlin * Oracle * Python * Security * Splunk * microservices * IBM * Network More: We are J.P. Morgan Asset Managements Trading Technology group, part of JPMorgan Chase, and we are modernizing our trading technology stack through a multi-year convergence journey. This Lead Site Reliability Engineer role is embedded within our software engineering team and focuses on front-office trading platforms in a fast-paced, global environment. We work closely with traders across Equities, Fixed Income, and FX, and collaborate with teams across EMEA, the US, and APAC. We value strong leadership, operational excellence, diversity and inclusion, and a first-class approach to serving our clients. This is a full-time role. ## Related Videos - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Monitoring as Code - Managing your dashboards at scale](https://www.wearedevelopers.com/videos/753-monitoring-as-code-managing-your-dashboards-at-scale) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Kotlin Multiplatform - True power of native code reuse](https://www.wearedevelopers.com/videos/4-kotlin-multiplatform-true-power-of-native-code-reuse) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Why Kotlin is the better Java and how you can start using it](https://www.wearedevelopers.com/videos/661-why-kotlin-is-the-better-java-and-how-you-can-start-using-it) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)