> Markdown version of [/jobs/ext/2624776-software-engineer-iii-devops-ml-ops-aws-streaming](https://www.wearedevelopers.com/jobs/ext/2624776-software-engineer-iii-devops-ml-ops-aws-streaming). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer III - DevOps/ML Ops AWS Streaming - **Company:** JPMorgan Chase & Co. - **Location:** Columbus, OH, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Automation of Tests, Unit Testing, Bash Shell, Code Generation, Software Quality, Continuous Delivery, Continuous Integration, Linux, DevOps, Elasticsearch, Python (Programming Language), Machine Learning, Networking Basics, Octopus Deploy, Performance Tuning, Reliability Engineering, Software Tools, Prometheus, Secure Coding, Software Engineering, Data Streaming, Toolchain, Management of Software Versions, Datadog, Autoscaling, Grafana, AWS ECS, Kubernetes, Low Latency, Apache Flink, Apache Kafka, Machine Learning Operations, Terraform, Stream Processing, Splunk - **Published:** August 2, 2026 - **Apply:** https://jpmc.fa.oraclecloud.com/hcmUI/CandidateExperience/en/sites/CX_1001/requisitions/preview/210737009 ## About the Role * Formal training or certification on software engineering concepts and 3+ years applied experience * Proven experience in DevOps, platform engineering, site reliability engineering, or machine learning operations supporting production services. * Hands-on experience with Amazon Web Services and Kubernetes (Amazon Elastic Kubernetes Service), with ability to provision, run, and troubleshoot production clusters. * Strong experience with Terraform, including modules, multi-environment patterns, remote state, and continuous integration/continuous delivery integration to reduce drift. * Experience operating Kafka-based streaming systems and resolving production pipeline issues related to performance, reliability, and resiliency. * Experience deploying and operating distributed compute platforms on Kubernetes, including Apache Flink. * Strong Linux and networking fundamentals with scripting and automation skills (Python and/or Bash) and a clear operational ownership mindset. * Hands-on experience using enterprise-authorized AI-assisted software development tools within the work environment (e.g., for coding, test creation, troubleshooting, or documentation) with demonstrated ability to critically evaluate, validate, and refine AI-generated outputs for correctness, performance, and security. * Understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; ability to guide peers on safe and effective usage within team practices., * Experience with GitOps and Kubernetes packaging tooling (for example, Argo CD, Flux, Helm, or Kustomize). * Experience with observability stacks such as Prometheus and Grafana, OpenTelemetry, Elasticsearch or OpenSearch, Splunk, or Datadog. * Experience with Kubernetes policy and security controls (for example, Open Policy Agent, Gatekeeper, or Kyverno) and secrets tooling (for example, Vault or external secrets controllers). * Experience with high-sensitivity domains such as authentication, identity verification, fraud, risk, or other regulated financial workloads. * Performance tuning experience for low-latency inference services and real-time streaming workloads. ## Description * Provision and manage Amazon Web Services Kubernetes and container environments, including networking, access controls, cluster configuration, and autoscaling, using Terraform to support consistent, repeatable delivery. * Build reusable infrastructure-as-code modules, define standards, and reduce configuration drift through strong environment hygiene and remote state management practices. * Enable end-to-end machine learning operations workflows across build, validation, packaging, deployment, monitoring, and retraining to support production model lifecycle needs. * Implement robust deployment and release patterns for machine learning-enabled services, including versioning, progressive delivery, and rollback strategies aligned to reliability goals. * Operate and troubleshoot Apache Kafka streaming integrations, addressing throughput, latency, resiliency, and consumer lag to sustain real-time decisioning workloads. * Deploy, operate, and scale Apache Flink on Kubernetes, managing job lifecycle, state, checkpoints, upgrades, recovery, and performance tuning. * Deploy, maintain, and scale Ray Serve on Kubernetes to provide low-latency inference and service orchestration, improving resource efficiency and runtime stability. * Implement and continuously improve observability (logs, metrics, tracing), dashboards, and alerting, and lead incident response and corrective actions to reduce mean time to recovery. * Leverages enterprise-authorized AI coding assist tools within the work environment to improve code quality, delivery speed, and productivity across complex deliverables (e.g., code generation/refactoring, unit test creation, documentation), while validating outputs through peer review, automated testing, and secure coding standards; contributes learnings and reusable patterns to improve broader team effectiveness. * Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation. ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [How Much FAANG Companies Actually Pay Software Engineers in 2025](https://www.wearedevelopers.com/magazine/230-how-much-faang-companies-actually-pay-software-engineers-in-2025) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)