> Markdown version of [/jobs/ext/624919-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/624919-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer - **Company:** Paramount Pictures - **Location:** Burbank, CA, United States - **Experience:** Expert - **Salary:** $130,200.0 - $195,300.0 - **Contract:** Temporary contract - **Skills:** Batch Processing, C++ (Programming Language), Nvidia CUDA, Continuous Integration, Distributed Computing Environment, Distributed Systems, Memory Management, Python (Programming Language), Machine Learning, Reliability Engineering, Prometheus, Azure Machine Learning, Reinforcement Learning, Pulumi, Graphics Processing Unit (GPU), Grafana, Multi-Cloud, HybridCloud, Kubernetes, Hardware Acceleration, Machine Learning Operations, Terraform, Software Version Control - **Published:** June 16, 2026 - **Apply:** https://careers.paramount.com/job/Burbank-Machine-Learning-Engineer-CA-91505/1387457000/ ## About the Role * 6-8+ years of experience in ML Infrastructure, Platform Engineering, or high-scale Backend Engineering. * Orchestration & Serving: Extensive experience with Kubernetes (K8s) and serving frameworks for large-scale ML models. * Hardware Proficiency: Strong knowledge of GPU architecture, CUDA, and optimizing ML workloads for hardware acceleration. * Leadership (IC4/5): Proven track record of owning the technical direction for a major domain anddriving impact across multiple teams., * Experience with Infra-as-Code (Terraform/Pulumi) and building automated MLOps pipelines. * Distributed Systems Mastery: Deep expertise with Ray (AnyScale) or similar distributed compute frameworks. * Familiarity with ML observability tools (Prometheus, Grafana, Weights & Biases, or MLFlow). * Experience managing multi-cloud or hybrid-cloud ML environments. * Deep knowledge of Python and C++ for performance-critical systems. ## Description We are seeking a Senior Lead / Lead ML Platform Engineer to architect and own the technical direction for our Training and Inference infrastructure. This is a high-leverage role designed for an expert who understands the deep technical stack required to shift ML models from research to global production. You will be responsible for the "engine room" of the AMLG, ensuring that our MLEs can train massive models efficiently and serve them with sub-millisecond reliability. This role requires a unique blend of expertise in distributed systems and hardware acceleration. You will lead the adoption and optimization of AnyScale (Ray) for distributed training and manage a high-performance Kubernetes-based inference environment. You aren't just managing clusters; you are building a seamless, scalable platform that abstracts the complexity of GPUs and distributed compute for the entire organization. Why This Role Matters The ML Platform Lead is the force-multiplier for every other ML pod. In this role, you will directly shape: * The Training Foundation: Establishing AnyScale/Ray as the standard for distributed compute, enabling MLEs to train models on petabytes of data without managing infrastructure. * Inference at Scale: Architecting the serving layer that handles billions of requests per day, optimizing for both p99 latency and GPU utilization. * Operational Excellence: Setting the organizational standards for how ML models are deployed, monitored, and scaled across the enterprise., * Technical Roadmap & Strategy: Own the long-term architectural direction for the Training and Inference domains, ensuring the platform scales 10x over a 1-3 year horizon. * Distributed Training Leadership: Lead the implementation and optimization of Ray/AnyScale, providing a unified compute layer for batch processing, model training, and reinforcement learning. * High-Performance Inference: Design and maintain K8s-based inference servers (e.g., Triton, TorchServe, or vLLM) optimized for GPU memory management and high throughput. * Hardware & Cost Optimization: Navigate the trade-offs between different GPU instances (A100s, H100s, T4s), optimizing for cost, availability, and performance. * Cross-Team Standardization: Solve high-leverage problems that affect multiple pods (e.g., Entry, Session, Presentation), establishing reusable patterns for CI/CD, model versioning, and canary deployments. * Reliability Engineering: Define and enforce SLIs/SLOs for the platform, ensuring that infrastructure failures never interrupt the user-facing personalization experience. * Mentorship & Coaching: Act as a technical mentor to senior engineers across the ML Platform and Applied ML pods, raising the bar for system design and operational rigor., * Unify the Compute Layer: Successfully transition the majority of AMLG training workloads to a governed AnyScale/Ray environment. * Optimize Inference ROI: Measurably improve GPU utilization and reduce inference costs through better auto-scaling and server optimization. * Establish Durable Standards: Author the "Gold Standard" for ML deployments that is adopted by at least three other pods in the organization. * Reduce Systemic Risk: Implement a self-healing infrastructure layer that significantly reduces manual intervention for cluster-related failures. ## Related Videos - [Why segmenting your infrastructure into tiers makes your infrastructure design better](https://www.wearedevelopers.com/videos/1960-why-segmenting-your-infrastructure-into-tiers-makes-your-infrastructure-design-better) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Terraform for Developers](https://www.wearedevelopers.com/videos/3-terraform-for-developers) - [The state of MLOps - machine learning in production at enterprise scale](https://www.wearedevelopers.com/videos/369-the-state-of-mlops-machine-learning-in-production-at-enterprise-scale) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)