Site Reliability Engineer -Jersey City, NJ & Dallas, TX

JPMorgan Chase & Co.
Jersey City, NJ, United States
1 day ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Compensation
$137,750.0 - $185,000.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) Adobe InDesign Artificial Intelligence Batch Processing Computer Programming Disaster Recovery Distributed Systems Middleware Fault Tolerance Monitoring of Systems Python (Programming Language) Log Analysis
+12 more
Windows PowerShell Release Management Reliability Engineering Site Reliability Engineering Practices Software Engineering Scripting Performance Testing Real Time Systems Data Management ArcSight Event Correlation SDET Golang

Job description

The Application Support Engineering role advances Site Reliability Engineering (SRE) practices for applications running in production. The role scope includes Application Support Engineers, SDETs, Software Engineers, and SREs focused on improving reliability, observability, recovery, automation, and operational readiness. This role brings production support expertise and engineering discipline earlier in the lifecycle to influence design, validate reliability requirements, reduce production risk, and drive measurable operational improvement. Your Primary Responsibilities

  • Participate in design reviews, sprint zero, and delivery planning to define and validate reliability requirements, including resiliency, observability, fault tolerance, performance, scalability, holiday and special-day processing, and disaster recovery.
  • Collaborate with Major Release Management to ensure each release meets SRE standards for observability, resiliency, and reliability requirements, support readiness, and knowledge base coverage.
  • Define and improve monitoring, observability, dashboards, telemetry coverage, and alert strategy to strengthen outage detection, reduce noise, improve signal quality, and accelerate incident response.
  • Assist in major incident response and root cause analysis by identifying observability gaps, improving telemetry and knowledge articles, and driving actions that reduce repeat incidents.
  • Drive automation, intelligent tooling, and AI-assisted remediation to reduce manual toil, improve consistency, accelerate recovery, and scale operational support.
  • Serve as the operational readiness authority before production releases by validating reliability requirements, assessing support readiness, surfacing production risks, and confirming release supportability.
  • Lead capacity, performance, workload trend, and resiliency analysis to ensure applications scale reliably under normal, peak, and stress conditions.
  • Establish and track reliability metrics such as availability, incident volume, MTTx, alert quality, automation coverage, reliability requirement compliance, change failure rate, and repeat incident reduction.
  • Participate in application reliability governance and service reviews by presenting incident trends, compliance metrics, operational risks, improvement actions, and readiness gaps.
  • Prepare executive reporting on reliability posture, release readiness, observability maturity, alert quality, incident trends, automation progress, risks, and improvement outcomes.
  • Promote SRE practices through mentoring, standards adoption, best-practice sharing, and approved AI tools that improve knowledge, observability, performance, security, and maintainability.

Requirements

  • Minimum of 10+ years of related technical experience across application support engineering, software engineering, site reliability engineering, production support, or application operations.
  • Bachelor’s degree preferred or equivalent practical experience.
  • Experience supporting business-critical applications in production environments.
  • SRE, observability, automation, or ITIL certifications are a plus.

Talents Needed for Success

  • Proven experience in one or more in-scope roles including Application Support Engineer, SDET, Software Engineer, or SRE, with responsibility for improving reliability practices, application validation, observability coverage, automation frameworks, and reliability standards.
  • Strong understanding of monitoring and observability platforms, including dashboard design, alert tuning, telemetry coverage, log analysis, metrics, traces, and event correlation.
  • Programming or scripting proficiency in one or more languages such as Python, Java, Go, PowerShell, or similar for automation, tooling, and operational efficiency.
  • Familiarity with distributed applications, middleware, messaging, batch processing, real-time processing, and production application behavior in high-availability environments.
  • Experience in financial services, capital markets, regulated environments, or other high-availability operational settings.
  • Demonstrated participation in disaster recovery, performance testing, resiliency testing, release readiness, incident response, and root cause analysis.
  • Knowledge of AI concepts, data platforms, anomaly detection, incident correlation, and intelligent automation use cases.
  • Strong collaboration skills across application support, application development, release management, risk, security, business, and vendor stakeholders.
  • Ability to translate production support insights into actionable engineering improvements that reduce risk, improve stability, and enhance customer experience.

Benefits & conditions

  • $137,750-185,000 per year There’s nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world’s most compl…

  • 5 days ago + *

Lead Site Reliability Engineer - Operations Excellence for AI Platforms JPMorgan Chase

  • Jersey City, NJ
  • $156,750-215,000 per year Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As…

  • 5 days ago + *

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

1:04 min

Introduction to Bitcoin script parsing tools

Steve Shadders · LIVE

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

1:06 min

Developer experience and project variety at scale

Alexandra Petri · World Congress 2023

1:53 min

Evaluating traditional scripting languages for modern development tasks

Jens Knipper Jens Knipper · Europe 2026 Virtual

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

Videos

See all

Related articles

See all