> Markdown version of [/jobs/ext/2967887-lead-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2967887-lead-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Site Reliability Engineer - **Company:** JPMorgan Chase & Co. - **Location:** Jersey City, NJ, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Bash Shell, Business Process Management, Continuous Integration, Python (Programming Language), Reliability Engineering, Large Language Models, Mttr, Terraform - **Published:** September 17, 2026 - **Apply:** https://www.jobmonkeyjobs.com/career/28026052/Lead-Site-Reliability-Engineer-New-Jersey-Jersey-City-7463 ## About the Role * 5+ years in SRE / production engineering / platform reliability / infrastructure operations at enterprise scale * Demonstrated success driving cross-team standardization and measurable reliability outcomes through influence and operating mechanisms. * Deep knowledge of SLOs/SLIs, error budgets, observability, incident response, postmortems, and change reliability. * Strong experience partnering with Engineering and Product leadership to align priorities and deliver results. * Hands-on experience with AWS and/or Azure and automation/IaC fundamentals (e.g., Python/Go/Bash, Terraform, CI/CD). * Experience using AI/LLM-enabled approaches in operations (AIOps, AI-assisted troubleshooting, agentic workflows) with appropriate controls and measurement., * Improved SLO attainment and reduced Sev1/Sev2 frequency; fewer repeat incidents. * Reduced MTTR/MTTI and lower change failure rate; reduced customer minutes impacted (cost of failure). * Measurable ticket reduction/deflection via scaled automation and self-service adoption. * Consistent operating model adopted across SRE teams with trusted executive reporting. ## Description As a Lead Site Reliability Engineer at JPMorgan Chase within the Public Cloud team, you will blend hands-on engineering with program leadership to promote platform stability, ensure consistent execution across SRE teams, and partner closely with Engineering and Product to deliver measurable improvements in availability, support outcomes, and cost of failure., * Drive consistency across SRE teams: establish and scale a "gold standard" for SLOs, on-call readiness, runbooks, postmortems, action tracking, and operational readiness across AWS/Azure/GCP. * Partner with Engineering and Product: embed reliability outcomes into roadmaps and delivery plans; convert incidents, support signals, and error budget trends into prioritized backlog with measurable impact. * Operational excellence & BPMs: design and run operating cadences (incident/stability reviews, KPI reviews, planning inputs), standardize intake/prioritization, and ensure closed-loop execution. * Risk governance: align reliability operations to risk/control expectations; operationalize recurring remediations (e.g., configuration drift, repeat findings) through centralized automation without sacrificing velocity. * Ticket reduction & automation opportunities: identify top drivers of support load and operational toil; build and maintain an automation opportunity pipeline; track adoption and deflection. * Platform stability metrics: own cross-platform reporting (SLO attainment, incident trends, MTTR/MTTI, change failure rate, ticket deflection, customer impact/cost of failure). * AI for reliability operations: apply AI/LLMs to improve triage, incident summarization, correlation/symptom mapping, and guardrailed automation; measure accuracy, safety, and outcomes. * Hands-on leadership: stay close to designs and critical implementations; lead systemic remediation and major incident response improvements. ## Related Videos - [Reducing Cognitive Overload Through Platform Engineering](https://www.wearedevelopers.com/videos/679-reducing-cognitive-overload-through-platform-engineering) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Old tools, new tricks](https://www.wearedevelopers.com/videos/1916-old-tools-new-tricks) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Your AI Agent is just a while loop with an API call. Let me prove it](https://www.wearedevelopers.com/videos/1988-your-ai-agent-is-just-a-while-loop-with-an-api-call-let-me-prove-it) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)