Senior Site Reliability Engineer/Platform Engineer - MT

JPMorgan Chase & Co.
Wilmington, DE, United States
about 1 month ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$152,000.0 - $215,000.0
Working hours
Regular working hours

Tech stack

JavaScript (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Backup Devices Continuous Integration Programming Tools Disaster Recovery Python (Programming Language) Reliability Engineering Prometheus Software Engineering Systems Integration
+9 more
TypeScript Datadog Grafana Reliability of Systems Infrastructure Automation Frameworks Deployment Automation Splunk New Relic (SaaS) Golang

Job description

We are looking for a highly hands-on Senior Site Reliability Engineer/Platform Engineer to improve the reliability, scalability, and operational maturity of our technology environment. This is not solely an infrastructure administration role. The person will combine strong systems and cloud/platform expertise with software engineering skills to design, build, and automate the services that support our engineering organization. A key part of the role will be connecting and enhancing observability systems across the environment-building integrations, automation, dashboards, alerting workflows, and reliability tooling that give teams actionable visibility into system health and performance. You will partner closely with engineering and infrastructure stakeholders to strengthen platform reliability, reduce operational friction, improve incident response, and establish scalable foundations for future product and AI initiatives., Build, maintain, and automate platform capabilities that improve system reliability, scalability, and developer productivity. Develop code, scripts, integrations, and internal tooling to connect observability, monitoring, alerting, and incident-management systems. Design and evolve observability practices across logs, metrics, traces, dashboards, alerting, and service health reporting. Improve CI/CD, deployment automation, environment consistency, and operational workflows through Infrastructure as Code and automation. Own reliability-focused initiatives including incident response, root-cause analysis, capacity planning, disaster recovery, backup strategy, and service resilience. Partner with software engineers to establish SRE standards, production readiness practices, and service-level objectives. Support and modernize core infrastructure, including cloud, virtualized, networked, and on-premise environments where applicable. Identify repetitive operational work and proactively replace it with scalable, maintainable automation.

Requirements

Strong software engineering experience, ideally with Python, Go, JavaScript/TypeScript, or a similar language used for automation and integrations. Deep experience with cloud/platform engineering, Infrastructure as Code, CI/CD, containers, and production operations. Hands-on experience with observability tooling such as Datadog, Grafana, Prometheus, ELK/OpenSearch, New Relic, Splunk, or similar platforms. Experience building monitoring integrations, alerting workflows, dashboards, operational APIs, or internal developer tools. Strong understanding of SRE principles, incident management, reliability engineering, and operational best practices. Ability to operate independently, set technical direction, and work across both infrastructure and software engineering teams. 3-6 month engagement (possibility extension) Location: Mexico and Colombia

Benefits & conditions

  • $152,000-215,000 per year If you are excited about shaping the future of technology and driving significant business impact in financial services, we are looking for people just like you. Join our team and …

  • 5 days ago + *

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all