> Markdown version of [/jobs/ext/615171-machine-learning-operations-engineer](https://www.wearedevelopers.com/jobs/ext/615171-machine-learning-operations-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Operations Engineer - **Company:** Garner Health - **Location:** New York, NY, United States - **Experience:** Expert - **Salary:** $256,000.0 - $285,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Software as a Service, Cloud Computing, Continuous Integration, Data Infrastructure, Monitoring of Systems, Machine Learning, Azure Machine Learning, Software Engineering, Containerization, Kubernetes, Machine Learning Operations, Terraform - **Published:** June 23, 2026 - **Apply:** https://www.dice.com/job-detail/25425700-fd04-4dd4-83cb-0e1883847791 ## About the Role * 5+ years of software engineering experience, with meaningful time spent operating ML or data-intensive systems in production. * Hands-on experience with the modern ML production stack: model serving (e.g., Sagemaker, Triton, or equivalent), feature stores, model registries, and CI/CD for ML. * Strong infrastructure and platform engineering fundamentals: Kubernetes, containerization, cloud (AWS preferred), Terraform/IaC, observability, and incident response. * Experience building ML platforms or significant components of one (not strictly consuming SaaS), with sound judgment around when to build vs. buy. * Strong collaboration with ML, data, platform engineers, data scientists, and product engineering teams, with the ability to lead projects and influence technical decisions. * Healthcare, regulated-data, or other high-stakes production ML experience is a plus but not required. * A desire to be a part of a high-performing, mission-driven team that operates with intense urgency, a strong sense of individual accountability, and a commitment to authentic feedback ## Description We are seeking a Senior MLOps Engineer to join our Platform Engineering team. This role will report to the Platform Engineering Manager, Developer Experience. As an early member of Garner's MLOps function, you will help build and operate the production machine learning systems that power our products, partnering closely with our machine learning and data science teams to enable the secure and consistent deployment of models. Given that these models directly influence health outcomes and cost-effectiveness for millions of patients, maintaining the highest standards of production quality is imperative. Where you will work: This role will be based in our New York City office (in the Financial District). You must be willing to work in the office 3 days per week on Tuesday, Wednesday and Thursday. What you will do: * Help ensure the reliability, performance, functionality, and cost-efficiency of Garner's production ML systems, contributing to SLOs, observability, and on-call responsibilities. * Build key components of Garner's ML platform, including data infrastructure (such as a feature store, model registry, and CI/CD for models) and standardized service patterns. * Implement ML-specific CI/CD pipelines: Help transition our deployment process from manual notebook hand-offs to automated, PR-driven CI/CD workflows that include automated data quality checks and statistical model validation prior to deployment. * Drive down cost and latency through improved architecture, hardware choices, and model optimization as appropriate. * Contribute to the workflows, standards, and KPIs that support a growing MLOps function, helping teammates and stakeholders quickly identify the health of the team's products and focus on areas where issues reside. * Help establish drift monitoring: Design and implement automated data drift and concept drift monitoring systems that alert the team when models degrade, laying the groundwork for future Continuous Training (CT) architectures. ## Related Videos - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Green Cloud Computing](https://www.wearedevelopers.com/videos/592-green-cloud-computing) - [Effective Machine Learning - Managing Complexity with MLOps](https://www.wearedevelopers.com/videos/185-effective-machine-learning-managing-complexity-with-mlops) - [Leverage Cloud Computing Benefits with Serverless Multi-Cloud ML ](https://www.wearedevelopers.com/videos/78-leverage-cloud-computing-benefits-with-serverless-multi-cloud-ml) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)