> Markdown version of [/jobs/ext/3050903-ml-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/3050903-ml-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML infrastructure engineer - **Company:** The Runway - **Location:** New York, NY, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Command-Line Interface, Cloud Engineering, Java GUIs, Python (Programming Language), Cloud Services, Azure Machine Learning, Management of Software Versions, Graphics Processing Unit (GPU), Pytorch, Large Language Models, Model Validation, Kubernetes, Data Management, Machine Learning Operations, Data Pipelines - **Published:** September 24, 2026 - **Apply:** https://startup.jobs/member-of-technical-staff-ml-platform-runwayml-10167398 ## About the Role * 5+ years of experience building ML infrastructure or data platforms in production environments, with at least some of that time spent on evaluation, experimentation, or benchmarking systems * Strong Python and PyTorch, and hands-on experience running large batch GPU workloads on Kubernetes * Experience designing data pipelines and storage for large volumes of media or model outputs, with attention to versioning and reproducibility * Familiarity with experimental statistics: paired comparisons, confidence intervals, multiple-comparison pitfalls, inter-rater agreement * Comfort building internal tools end to end, from the command line to the browser * Ability to lead a broad technical area: gather requirements from many teams, set direction, make tradeoffs, and drive a roadmap without waiting to be told * Familiarity with the full model development lifecycle: data, training, evaluation, serving * Self-starter who can work embedded with research teams and move fast * Strong systems thinking and pragmatic approach to production reliability * Humility and open mindedness; at Runway we love to learn from one another Nice to have * Track record of building evaluation suites for generative models * Hands-on work with LLM- or VLM-as-judge pipelines * Experience with online experimentation platforms and connecting offline metrics to product outcomes * Prior work evaluating agents or robotics policies ## Description We're looking for an ML infrastructure engineer to own model evaluation at Runway, end to end. Every decision we make about a model - which checkpoint to keep training, what to ship to millions of users, which datamix and architecture shows the most promise - rests on evals. Today that work is spread across the organization. You'll turn it into one platform. Your job will be to design and build the systems that generate samples at scale, score them with automated metrics and human annotations, track results across checkpoints and releases, and put the answers in front of researchers in minutes rather than days. You'll define what "better" means operationally - how we measure it, how confident we are, and how a result becomes a ship/no-ship decision. You'll be embedded within research teams as a member of ML Platform, and you'll set the technical direction for evals across the company. This is an extremely high-leverage role: the quality of our models is bounded by how well we can measure them. A peek at our technical stack Our inference and training code is written in Python (PyTorch) and is cloud-native. We leverage multiple clouds and use a combination of Kubernetes native and home-grown tooling for efficient job orchestration. We're big users of agentic development and operations. We have a robust internal platform and provide broad access to the latest models and harnesses. You'll have access to best-in-class models, agents, GPUs, storage and cloud services you need to drive a world class evaluation pipeline that helps us deliver frontier models. What you'll do * Own the evaluation platform end to end: the tooling and systems required to generate, annotate, review and adapt at frontier scale * Define the CLIs, APIs, GUIs and storage layers required to make this process seamless, fast, sophisticated and collaborative * Work directly with research teams on video, image, audio, agents, and robotics to understand what they need to measure and build it, then generalize the result into the platform * Help set the standards for how we evaluate models at Runway: reproducibility, metric definitions, reporting formats, and when a result is trustworthy enough to act on * Support broad adoption of the platform and its integration throughout Runway's research efforts across training, production model serving and more * Contribute broadly as a member of the ML Platform team to tools and systems that help Runway train and serve frontier models ## Related Videos - [Beam Me Up, Java! Unraveling the Warp-Speed Evolution: A Journey through Java LTS Versions 11 to 21](https://www.wearedevelopers.com/videos/658-beam-me-up-java-unraveling-the-warp-speed-evolution-a-journey-through-java-lts-versions-11-to-21) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [The state of MLOps - machine learning in production at enterprise scale](https://www.wearedevelopers.com/videos/369-the-state-of-mlops-machine-learning-in-production-at-enterprise-scale) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [ I Gave a Video Editor More Autonomy Than a Trading Bot. On Purpose.](https://www.wearedevelopers.com/magazine/773-i-gave-a-video-editor-more-autonomy-than-a-trading-bot-on-purpose) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)