Site Reliability Engineer III - AWS, Java and Kubernetes

JPMorgan Chase & Co.
Chicago, IL, United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Java (Programming Language) Artificial Intelligence Amazon Web Services Cloud Computing Computer Programming Databases Continuous Integration Python (Programming Language) Network Troubleshooting Reliability Engineering Site Reliability Engineering Practices Ansible
+15 more
Prometheus Datadog Cloud Platform System Computer Network Technologies Grafana HybridCloud Gitlab Containerization Kubernetes Data Management Terraform Splunk Dynatrace Docker Jenkins

Job description

Experteer Overview In this role you will strengthen the reliability and performance of mission-critical systems within JPMorgan Chase’s Corporate Technology team. You will work on scalable, observable architectures using code and cloud infrastructure, and you will help the team adopt SRE best practices. You’ll contribute to incident triage, post-incident analysis, and proactive improvements tied to SLOs, delivering stable, scalable platforms. This is a hands-on opportunity to influence how we design, deploy, and run reliable systems at scale. Compensation / Benefits * Guide and promote SRE best practices within the team and foster consensus on design decisions * Collaborate with software engineers to design, implement, and test deployment and reliability approaches via CI/CD pipelines * Leverage enterprise AI capabilities to accelerate incident triage, troubleshooting, and post-incident analysis while handling data securely * Implement infrastructure, configuration, and network as code for assigned applications and platforms * Work with stakeholders to resolve complex problems and proactively address issues using SLOs/SLIs * Improve availability, reliability, and scalability of applications in collaboration with partners * Identify roadblocks and explore new technologies to solve business problems * Apply SRE fundamentals to monitoring, incident response, capacity planning, and toil reduction Tasks * Formal training or certification in Site Reliability Engineering concepts and 3+ years of applied experience * Exposure to SRE practices for data management/migration platforms; familiarity with on-prem/public cloud components (Compute, Storage, Networks, Database) * Understanding of SRE fundamentals and ability to define and track SLOs/SLIs * Proficiency in at least one programming or configuration tool (e.g., Python, Ansible, Terraform) * Experience with observability and telemetry tools (e.g., Grafana, Dynatrace, Prometheus, Datadog, Splunk) * Experience with public/private/hybrid cloud environments and container orchestration (Kubernetes, ECS, Docker) * Experience with CI/CD tools (Jenkins, GitLab, Terraform) * Familiarity with troubleshooting networking technologies * Experience with containerization and orchestration and related networking issues Key requirements * competitive base salary * incentive compensation may be available * comprehensive health care coverage * retirement savings plan * backup childcare * tuition reimbursement (education assistance)

Requirements

for assigned applications and platforms * Work with stakeholders to resolve complex problems and proactively address issues using SLOs/SLIs * Improve availability, reliability, and scalability of applications in collaboration with partners * Identify roadblocks and explore new technologies to solve business problems * Apply SRE fundamentals to monitoring, incident response, capacity planning, and toil reduction Tasks * Formal training or certification in Site Reliability Engineering concepts and 3+ years of applied experience * Exposure to SRE practices for data management/migration platforms; familiarity with on-prem/public cloud components (Compute, Storage, Networks, Database) * Understanding of SRE fundamentals and ability to define and track SLOs/SLIs * Proficiency in at least one programming or configuration tool (e.g., Python, Ansible, Terraform) * Experience with observability and telemetry tools (e.g., Grafana, Dynatrace, Prometheus, Datadog, Splunk) * Experience with aaaaaaaaaaa Kubernetes cloud environments and container orchestration (Kubernetes, ECS, Docker) * Experience with CI/CD tools (Jenkins, GitLab, Terraform) * Familiarity with troubleshooting networking technologies * Experience with containerization and orchestration and related networking issues Key requirements * competitive base salary * incentive compensation may be available * comprehensive health care coverage * retirement savings plan * backup childcare * tuition reimbursement (education assistance)

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

6:14 min

Structuring CI/CD pipelines with integrated security and quality checks

Christoph Ruggenthaler · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

4:54 min

Implementing geographic salary tiers for compensation equity and fairness

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all