Lead Site Reliability Engineer

JPMorgan Chase & Co.
London, UK
23 days ago
Apply on www.adzuna.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
£62,000.0 - £102,000.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) Artificial Intelligence Algorithmic Trading Automation of Tests Software Bug Management Cloud Computing Cloud Engineering Cyber Security Continuous Integration Disaster Recovery Distributed Systems Python (Programming Language)
+18 more
Operational Data Store Oracle (Applications) Performance Tuning Systems Development Life Cycle Reliability Engineering Software Engineering Grafana Kotlin Event Driven Architecture Low Latency Influxdb Deployment Automation Production Code Apache Kafka Splunk Dynatrace Microservices Oracledb

Job description

  • Engage daily with traders across asset classes to understand workflows, pain points, and reliability priorities.
  • Act as a trusted engineering partner to the desk, ensuring systems are stable, performant, and aligned with business needs.
  • Support live trading environments, including incident response, root cause analysis, and post-mortem leadership.
  • Participate as a core member of the software engineering team in daily standups and design discussions.
  • Contribute directly to the codebase to implement reliability improvements, performance optimisations, bug fixes, and automation.
  • Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self-healing workflows, and resilience engineering.
  • Use enterprise-authorized AI capabilities to accelerate major-incident triage, troubleshooting, and post-incident analysis, while validating outputs and handling operational data securely.
  • Drive reuse-first adoption of AI-assisted reliability workflows across SDLC and toolchain practices, ensuring traceability, auditability, resiliency, and security controls.
  • Drive improvements in latency, throughput, and stability across high-volume trading applications.
  • Build and maintain tooling for monitoring, alerting, and distributed tracing across global environments.
  • Operate within a globally distributed engineering and trading organization, collaborating with teams in EMEA, US, and APAC.
  • Partner with infrastructure, networking, cloud engineering, and cybersecurity teams to ensure end-to-end reliability.

Technologies:

  • AI
  • CI/CD
  • Cloud
  • Dynatrace
  • Embedded
  • Grafana
  • Support
  • Java
  • Kafka
  • Kotlin
  • Oracle
  • Python
  • Security
  • Splunk
  • microservices
  • IBM
  • Network

More:

We are J.P. Morgan Asset Managements Trading Technology group, part of JPMorgan Chase, and we are modernizing our trading technology stack through a multi-year convergence journey. This Lead Site Reliability Engineer role is embedded within our software engineering team and focuses on front-office trading platforms in a fast-paced, global environment. We work closely with traders across Equities, Fixed Income, and FX, and collaborate with teams across EMEA, the US, and APAC. We value strong leadership, operational excellence, diversity and inclusion, and a first-class approach to serving our clients. This is a full-time role.

Requirements

  • Strong hands-on experience in front-office trading environments or similarly high-pressure, low-latency domains.
  • Proficiency with SRE tooling and techniques, including FIX messaging, Kafka, Grafana, Splunk, ITRS Geneos, Dynatrace, InfluxDB, MQ, and Oracle DB.
  • Demonstrated experience using enterprise-authorized AI capabilities to improve SRE workflows, with strong validation habits and awareness of data sensitivity.
  • Ability to evaluate AI-assisted operational recommendations for correctness and risk, define guardrails for team usage, and ensure outcomes align with resiliency and security expectations.
  • Deep knowledge of reliability engineering principles, including SLIs/SLOs, real-time telemetry, disaster recovery planning, capacity planning, and performance tuning.
  • Experience designing and implementing observability frameworks for mission-critical systems.
  • Proven ability to lead incident response and drive long-term remediation.
  • Solid programming skills in Python, Java, or Kotlin, with the ability to contribute production-grade code.
  • Experience with microservices, distributed systems, and event-driven architectures.
  • Strong understanding of CI/CD pipelines, automated testing, and deployment strategies.
  • Comfort interacting directly with traders and senior stakeholders.
  • Excellent communication skills, especially in translating technical issues into business impact.
  • Ability to operate calmly and decisively in high-pressure situations.
  • Strong leadership presence with a collaborative mindset.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.co.uk
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:46 min

Introduction to the speaker and engineering background

Llywelyn Griffith-Swain · World Congress 2023

1:07 min

Architecting the availability stack with Prometheus and Grafana

Gabriel Labachelerie · World Congress 2023

3:09 min

Understanding Kotlin Multiplatform and its compiler targets

Petar Marijanović · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:18 min

Recommended resources and frameworks for learning Kotlin natively

Iris Hunkeler · LIVE

Videos

See all

Related articles

See all