> Markdown version of [/jobs/ext/1663167-senior-machine-learning-engineer-ml-infrastructure-online](https://www.wearedevelopers.com/jobs/ext/1663167-senior-machine-learning-engineer-ml-infrastructure-online). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Machine Learning Engineer, ML Infrastructure- Online - **Company:** Unity Technologies - **Location:** Seattle, WA, United States - **Experience:** Expert - **Salary:** $187,200.0 - $243,300.0 - **Contract:** Permanent contract - **Skills:** Airflow, Computer Programming, Distributed Systems, Python (Programming Language), Machine Learning, Tensorflow, Azure Machine Learning, Pytorch, Autoscaling, Caching, Kubernetes, Low Latency, Deployment Automation, Machine Learning Operations - **Published:** July 25, 2026 - **Apply:** https://dejobs.org/x/x/741E3CFB9DB04C00A05FDA6592B312B9/job/ ## About the Role * Experience building and operating production-grade online ML inference systems, such asNVIDIA Triton Inference Server, TorchServe, Ray Serve, TensorFlow Serving, or similar systems. * Experience with model serving frameworks such as NVIDIA Triton Inference Server, TorchServe, Ray Serve, TensorFlow Serving, or similar systems. * Experience optimizing inference workloads using techniques such as dynamic batching, model compilation, quantization, GPU acceleration, GPU kernel optimization, caching, or runtime tuning. * Strong experience with distributed systems, Kubernetes, autoscaling, service reliability, and production observability. * Strong programming skills in Python, with practical experience working on production ML systems and high-scale services. * Experience with PyTorch and modern model deployment workflows, including model packaging, validation, and serving lifecycle management. * Experience designing infrastructure for safe model rollout, canary testing, A/B experimentation, and automated rollback. * Strong systems thinking, with the ability to reason about latency, throughput, reliability, scalability, and cost tradeoffs in online systems. * Proven ability to lead technical direction and influence architectural decisions across teams without formal authority. ## Description We are seeking a Senior ML engineer to design and evolve Unity Vector's online model inference platform. This role focuses on building reliable infrastructure for serving machine learning models in production, optimizing inference performance, and enabling safe, efficient experimentation across high-traffic online systems. You will work closely with ML engineers, platform teams, and product stakeholders to ensure models can be deployed, scaled, monitored, and iterated on efficiently. You will play a key role in shaping how models are packaged, served, validated, monitored, and optimized in production environments. This role requires strong systems thinking, deep experience with production ML infrastructure, and the ability to drive architectural improvements across teams. What you'll be doing * Design and operate large-scale online inference infrastructure that serves production ML models with low latency and high reliability, such as PyTorch, Triton Inference Server, Kubernetes, GKE, Ray, or similar distributed serving frameworks. * Develop infrastructure that supports distributed training workflows using technologies such as Pytorch, Ray Data, and Ray Train, etc. * Integrate ML pipelines with workfloworchestration systems (e.g., Flyte, Airflow, or similar)to enable reliable multi-stage training workflows * Optimize model performance throughmodel compilation, GPU/CPU utilization improvements, request scheduling, kernel fusion, and runtime-level tuning. * Improve observability of ML systems through latency, throughput, error-rate, cost, saturation, and model-health monitoring. * Partner closely with ML engineers to support faster model iteration while maintaining production safety, scalability, and cost efficiency. * Improve the reliability and reproducibility of model serving workflows, including model packaging, artifact validation, compatibility testing, and deployment automation. * Lead architectural improvements that make the online ML platform more robust, user-friendly, scalable, and cost-efficient., This position requires the incumbent to have a sufficient knowledge of English to have professional verbal and written exchanges in this language since the performance of the duties related to this position requires frequent and regular communication with colleagues and partners located worldwide and whose common language is English. This posting is intended to fill an existing vacancy, and we are committed to providing applicants with updates throughout the hiring process in accordance with applicable law. Headhunters and recruitment agencies may not submit resumes/CVs through this website or directly to managers. Unity does not accept unsolicited headhunter and agency resumes. Unity will not pay fees to any third-party agency or company that does not have a signed agreement with Unity. ## Related Videos - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How machine learning can help us tell fact from fiction](https://www.wearedevelopers.com/magazine/509-how-machine-learning-can-help-us-tell-fact-from-fiction) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)