> Markdown version of [/jobs/ext/1306197-lead-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1306197-lead-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Site Reliability Engineer - **Company:** JPMorgan Chase & Co. - **Location:** London, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Algorithmic Trading, Automation of Tests, Software Bug Management, Cloud Engineering, Cyber Security, Continuous Integration, Disaster Recovery, Distributed Systems, IBM WebSphere MQ, Python (Programming Language), Knowledge Management, Operational Data Store, Performance Tuning, Systems Development Life Cycle, Reliability Engineering, Software Engineering, Grafana, Kotlin, Event Driven Architecture, Low Latency, Influxdb, Deployment Automation, Apache Kafka, Splunk, Dynatrace, Microservices, Oracledb - **Published:** July 17, 2026 - **Apply:** https://jpmc.fa.oraclecloud.com/hcmUI/CandidateExperience/en/sites/CX_1001/requisitions/preview/210768948 ## About the Role Our trading technology stack is undergoing a multi-year convergence and modernization journey. You will play a pivotal role in shaping our next-generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast-paced front-office environments, enjoys direct interaction with traders, and wants to influence the reliability culture of a major global trading organization., * Strong hands on experience in front office trading environments or similarly high pressure, low latency domains. * Proficiency with SRE tooling and techniques, including FIX messaging, Kafka, Grafana, Splunk, ITRS Geneos, Dynatrace, InfluxDB, MQ (IBM MQ or similar), Oracle DB * Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity. * Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations. * Deep knowledge of reliability engineering principles: SLIs/SLOs, real-time telemetry, disaster recovery planning, capacity planning, and performance tuning. * Experience designing and implementing observability frameworks for mission critical systems. * Proven ability to lead incident response and drive long term remediation. * Solid programming skills in Python, Java, or Kotlin, with the ability to contribute production grade code. Experience with microservices, distributed systems, and event driven architectures. * Strong understanding of CI/CD pipelines, automated testing, and deployment strategies. * Comfortable interacting directly with traders and senior stakeholders. Excellent communication skills, especially when translating technical issues into business impact. Ability to operate calmly and decisively in high pressure situations. * Strong leadership presence with a collaborative mindset. ## Description * Engage daily with traders across asset classes (Equities, Fixed Income, FX) to understand workflows, pain points, and reliability priorities. * Act as a trusted engineering partner to the desk, ensuring systems are stable, performant, and aligned with business needs. * Support live trading environments, including incident response, root cause analysis, and post mortem leadership. Work as a core member of the software engineering team, participating in daily standups and design discussions. * Contribute directly to the codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation. * Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self healing workflows, and resilience engineering. * Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements. * Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls. * Drive improvements in latency, throughput, and stability across high volume trading applications. * Build and maintain tooling for monitoring, alerting, and distributed tracing across global environments. * Operate within a globally distributed engineering and trading organization, collaborating with teams in EMEA, US, and APAC. * Partner with infrastructure, networking, cloud engineering, and cybersecurity teams to ensure end to end reliability. ## Related Videos - [Inside Bitpanda's Tech Stack: Scaling a European Fintech Leader - Markus Dorner](https://www.wearedevelopers.com/videos/1979-inside-bitpanda-s-tech-stack-scaling-a-european-fintech-leader-markus-dorner) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Kotlin Multiplatform - True power of native code reuse](https://www.wearedevelopers.com/videos/4-kotlin-multiplatform-true-power-of-native-code-reuse) - [Microservices architecture as a key element in building trading systems for global finance markets](https://www.wearedevelopers.com/videos/1196-microservices-architecture-as-a-key-element-in-building-trading-systems-for-global-finance-markets) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Best Companies to work for in London: Top 25 Companies in 2023](https://www.wearedevelopers.com/magazine/187-best-companies-to-work-for-in-london-top-25-companies-in-2023)