Software Engineer III - DevOps/ML Ops AWS Streaming

JPMorgan Chase & Co.
Columbus, OH, United States
about 1 month ago
Apply on jpmc.fa.oraclecloud.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Automation of Tests Unit Testing Bash Shell Code Generation Software Quality Continuous Delivery Continuous Integration Linux DevOps Elasticsearch
+25 more
Python (Programming Language) Machine Learning Networking Basics Octopus Deploy Performance Tuning Reliability Engineering Software Tools Prometheus Secure Coding Software Engineering Data Streaming Toolchain Management of Software Versions Datadog Autoscaling Grafana AWS ECS Kubernetes Low Latency Apache Flink Apache Kafka Machine Learning Operations Terraform Stream Processing Splunk

Job description

  • Provision and manage Amazon Web Services Kubernetes and container environments, including networking, access controls, cluster configuration, and autoscaling, using Terraform to support consistent, repeatable delivery.
  • Build reusable infrastructure-as-code modules, define standards, and reduce configuration drift through strong environment hygiene and remote state management practices.
  • Enable end-to-end machine learning operations workflows across build, validation, packaging, deployment, monitoring, and retraining to support production model lifecycle needs.
  • Implement robust deployment and release patterns for machine learning-enabled services, including versioning, progressive delivery, and rollback strategies aligned to reliability goals.
  • Operate and troubleshoot Apache Kafka streaming integrations, addressing throughput, latency, resiliency, and consumer lag to sustain real-time decisioning workloads.
  • Deploy, operate, and scale Apache Flink on Kubernetes, managing job lifecycle, state, checkpoints, upgrades, recovery, and performance tuning.
  • Deploy, maintain, and scale Ray Serve on Kubernetes to provide low-latency inference and service orchestration, improving resource efficiency and runtime stability.
  • Implement and continuously improve observability (logs, metrics, tracing), dashboards, and alerting, and lead incident response and corrective actions to reduce mean time to recovery.
  • Leverages enterprise-authorized AI coding assist tools within the work environment to improve code quality, delivery speed, and productivity across complex deliverables (e.g., code generation/refactoring, unit test creation, documentation), while validating outputs through peer review, automated testing, and secure coding standards; contributes learnings and reusable patterns to improve broader team effectiveness.
  • Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.

Requirements

  • Formal training or certification on software engineering concepts and 3+ years applied experience
  • Proven experience in DevOps, platform engineering, site reliability engineering, or machine learning operations supporting production services.
  • Hands-on experience with Amazon Web Services and Kubernetes (Amazon Elastic Kubernetes Service), with ability to provision, run, and troubleshoot production clusters.
  • Strong experience with Terraform, including modules, multi-environment patterns, remote state, and continuous integration/continuous delivery integration to reduce drift.
  • Experience operating Kafka-based streaming systems and resolving production pipeline issues related to performance, reliability, and resiliency.
  • Experience deploying and operating distributed compute platforms on Kubernetes, including Apache Flink.
  • Strong Linux and networking fundamentals with scripting and automation skills (Python and/or Bash) and a clear operational ownership mindset.
  • Hands-on experience using enterprise-authorized AI-assisted software development tools within the work environment (e.g., for coding, test creation, troubleshooting, or documentation) with demonstrated ability to critically evaluate, validate, and refine AI-generated outputs for correctness, performance, and security.
  • Understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; ability to guide peers on safe and effective usage within team practices., * Experience with GitOps and Kubernetes packaging tooling (for example, Argo CD, Flux, Helm, or Kustomize).
  • Experience with observability stacks such as Prometheus and Grafana, OpenTelemetry, Elasticsearch or OpenSearch, Splunk, or Datadog.
  • Experience with Kubernetes policy and security controls (for example, Open Policy Agent, Gatekeeper, or Kyverno) and secrets tooling (for example, Vault or external secrets controllers).
  • Experience with high-sensitivity domains such as authentication, identity verification, fraud, risk, or other regulated financial workloads.
  • Performance tuning experience for low-latency inference services and real-time streaming workloads.

Benefits & conditions

We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process.

About the company

Chase is a leading financial services firm, helping nearly half of America’s households and small businesses achieve their financial goals through a broad range of financial products. Our mission is to create engaged, lifelong relationships and put our customers at the heart of everything we do. We also help small businesses, nonprofits and cities grow, delivering solutions to solve all their financial needs.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jpmc.fa.oraclecloud.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:09 min

Evaluating mature stream processing frameworks for production systems

Soroosh Khodami Soroosh Khodami · World Congress 2024

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all