> Markdown version of [/jobs/ext/1728050-senior-machine-learning-operations-engineer](https://www.wearedevelopers.com/jobs/ext/1728050-senior-machine-learning-operations-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Machine Learning Operations Engineer - **Company:** Paramount Global - **Location:** New York, NY, United States - **Experience:** Expert - **Salary:** $139,200.0 - $208,800.0 - **Contract:** Permanent contract - **Skills:** Information Engineering, DevOps, Monitoring of Systems, Machine Learning, Standard Sql, Management of Software Versions, Delivery Pipeline, Machine Learning Operations - **Published:** July 14, 2026 - **Apply:** https://www.manhattanjobs.com/job.asp?id=3318684093&tx=KL3737FFF&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role * 5+ years in ML engineering, applied ML, or a related ML role, with demonstrated experience on the operational side of monitoring, reliability, deployment, or incident response * Has built or operated model registries, ML monitoring systems, or production ML pipelines * Understands ML systems end-to-end not just the infra layer, but why a stale feature or a shifted distribution matters * Robust SQL skills and comfort digging into data distributions, feature health, and model behavior * Comfortable partnering with DevOps and Platform teams to define infrastructure needs without needing to own the infra yourself ## Description We're hiring a Senior Machine Learning Operations Engineer to own the operational layer around our personalization and recommendation Machine Learning (ML) systems. Our models retrain and deploy daily on automated pipelines. Your job is to make sure we can trust what's running, know when something is off, and fix it fast. You'll sit within DevOps and work closely with ML engineers, who own the models end-to-end. You won't be building infrastructure from scratch, you'll partner with DevOps and Platform Engineering to get the tooling you need, then own it day-to-day. What You'll Do * Own model traceability: Every model in production should have clear lineage: what data trained it, what code produced it, what validation it passed, and how it's performing. Evaluate and recommend tooling for versioning, metadata, and model registry, and work with MLEs to drive adoption. * Build end-to-end monitoring: Monitor the full signal path: data arrival, feature distribution stability, model metrics, and serving latency against SLA. Own this individually, don't rely solely on upstream teams to catch their own issues. * Partner with Data Engineering on data quality: Collaborate to surface data quality issues, detect drift in upstream sources, and ensure features stay fresh and reliable. * Detect issues proactively: Track drift over weeks, flag slow degradation before it crosses a threshold, surface feature freshness problems before they cascade. * Build diagnostic tooling: When something goes wrong, get from "recommendations look off" to root cause in minutes. That means ensuring the right context is logged at each stage, candidates, features, serving context, and building the dashboards to tie it collectively. * Own incident response for ML systems: Maintain rollback playbooks and pre-defined hotfix strategies with quantified tradeoffs. Own automated gates that block bad deployments. Run post-mortems and close the gaps. * Coordinate on post-deployment metrics: Work with ML engineers, data engineers, and stakeholders to define what metrics to collect after deployment and why they matter. ## Related Videos - [DevOps for Machine Learning](https://www.wearedevelopers.com/videos/179-devops-for-machine-learning) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [How building an industry DBMS differs from building a research one](https://www.wearedevelopers.com/videos/768-how-building-an-industry-dbms-differs-from-building-a-research-one) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Effective Machine Learning - Managing Complexity with MLOps](https://www.wearedevelopers.com/videos/185-effective-machine-learning-managing-complexity-with-mlops) - [How Machine Learning is turning the Automotive Industry upside down](https://www.wearedevelopers.com/videos/61-how-machine-learning-is-turning-the-automotive-industry-upside-down) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How machine learning can help us tell fact from fiction](https://www.wearedevelopers.com/magazine/509-how-machine-learning-can-help-us-tell-fact-from-fiction) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)