> Markdown version of [/jobs/ext/2603765-ml-platform-engineer](https://www.wearedevelopers.com/jobs/ext/2603765-ml-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Platform Engineer - **Company:** Hadrian Inc. - **Location:** Los Angeles, CA, United States - **Salary:** $170,000.0 - $300,000.0 - **Contract:** Permanent contract - **Skills:** A/B Testing, Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, Data Analysis, Automation of Tests, Encodings, Continuous Integration, Distributed Systems, Python (Programming Language), Open Source Technology, Standard Sql, SQL Databases, Scripting, Graphics Processing Unit (GPU), Robot Operating System, Autoscaling, Fastapi, AI Platforms, Kubernetes, Machine Learning Operations, Api Design - **Published:** August 6, 2026 - **Apply:** https://www.careerbuilder.com/job-details/ml-platform-engineer-los-angeles-ca--1038f1d2-a2fe-4dbd-b143-e50433eaf523 ## About the Role * Track record building and operating production ML infrastructure across multiple models or inference workloads. * Strong production-level Python and SQL skills, including typing, testing, packaging, API design, and building observability features. * Hands-on experience with Kubernetes, containers, and handling distributed-system failure modes such as retries, partial failures, idempotence, and resource isolation. * Engineering background with model registries, feature systems, batch/real-time inference, experiment tracking, or model CI/CD workflows. * Practical judgment around latency, throughput, availability, multi-tenancy, autoscaling, and infrastructure cost optimizations. * Ability to build stable interfaces and collaborate closely with engineering and scientific stakeholders. What Will Set You Apart * Experience implementing feature stores (Feast, Tecton, or internal systems). * Production work with Ray Serve, KServe, Triton, BentoML, SageMaker, Vertex AI, or custom gRPC inference services. * Experience serving and evaluating vision, document-understanding, embedding, or generative pipelines. * Expertise in GPU inference optimization, multi-model serving, edge inference, or Go/Rust performance-sensitive AI services. * Background in regulated environments or open-source contributions to ML infrastructure projects (MLflow, Feast, KServe, Ray)., A/B Testing, Additive Manufacturing, Aerospace and Defense, Application Programming Interface (API), Artificial Intelligence (AI), Autoscaling, Casting, Contract Management, Cost Control, Data Analysis, Data Modeling, Data Science, Distributed Computing, Driver's License, Electronics Manufacturing, Federal Laws and Regulations, Forecasting, GPU (Graphics Processing Unit), Genetics, Government, Incident Response, Life Insurance, Machine Tool, Manufacturing, Medical Conditions, Military, Open Source, Operations Research, Performance Metrics, Python Programming/Scripting Language, Regulations, Resource Management, Robotics Software, SQL (Structured Query Language), Team Building, Team Player, Telemetry, Test Automation, Testing, Typing, United States Citizen, Vision Plan, Welding ## Description This is an ML infrastructure role at the core of Hadrian's technology stack. While Data Science, Operations Research, Vision, and Document AI teams build models, you will own the platform that ensures these models remain reliable, effective, and secure in production. You'll standardize our deployment patterns built around MLflow, Dagster, ECR, FastAPI, and EKS, making them the backbone for packaging, evaluating, releasing, serving, monitoring, and rolling back models across Hadrian's automated factories. What You'll Do * Build the production platform that enables Hadrian's factories to safely depend on models for drawing extraction, cycle-time prediction, forecasting, scheduling, and more-with measurable performance and fast rollback. * Develop shared batch and online serving for tabular, vision, document-AI, scheduling, graph, and embedding workloads, targeting clear SLAs for latency, availability, and isolation. * Create repeatable release and evaluation processes featuring automated tests, reproducible artifacts, lineage, shadow deployments, canaries, and A/B tests. * Own online feature serving and maintain contract integrity with offline feature tables; proactively detect and address training-serving skew, feature drift, bad data, and model degradation. * Build operational tooling for telemetry, incident response, autoscaling, resource and GPU management, cost attribution, and secure model routing. * Develop APIs, SDKs, reusable templates, and documentation that teams can adopt without requiring close support from platform engineers. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Intro to FastAPI](https://www.wearedevelopers.com/videos/462-intro-to-fastapi) - [JavaScript? No. Java Scripts! - Scripting with Java](https://www.wearedevelopers.com/videos/2094-javascript-no-java-scripts-scripting-with-java) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)