> Markdown version of [/jobs/ext/3110198-staff-ml-platform-engineer](https://www.wearedevelopers.com/jobs/ext/3110198-staff-ml-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff ML Platform Engineer - **Company:** Datavant - **Location:** Boston, MA, United States - **Experience:** Experienced - **Salary:** $224,000.0 - $280,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, Computer Vision, Big Data, Continuous Integration, Cursor (Graphical User Interface Elements), Github, Identity and Access Management, Java Virtual Machine (JVM), Python (Programming Language), Machine Learning, Open Source Technology, Tensorflow, Azure Machine Learning, Software Engineering, Pytorch, Large Language Models, Snowflake, Apache Spark, Caching, Kubernetes, Low Latency, Apache Kafka, Machine Learning Operations, Terraform, Databricks - **Published:** September 27, 2026 - **Apply:** https://dejobs.org/x/x/ADC6096379624E279FFD30157B203F5B/job/ ## About the Role * 10+ years of software engineering experience, with 3+ years designing, evolving, and operating enterprise-scale ML platforms in production * Strong technical judgment under ambiguity and a track record of setting standards, influencing peers, and raising the bar across teams * Hands-on production experience with Databricks and/or Amazon SageMaker, MLflow (or an equivalent tracking + registry system), and at least one core ML framework (PyTorch, TensorFlow, or similar) * Fluency in Java (or a JVM equivalent) and Python, with real depth in Apache Spark for large-scale data and distributed compute * Real depth in AWS: networking, IAM, GPU compute, and the storage and messaging services this role touches, with the judgment to know what to reach for and when * Fluency with Terraform, containers, Kubernetes, and GitHub-based CI/CD for ML workloads * Direct experience serving LLMs in production, including cost management, evaluation harnesses, and safe handling of sensitive prompts and outputs * AI-native working style: daily use of Claude Code, Cursor, Copilot, or equivalent, with views on how they make a team faster * Clear written and verbal communication, especially in async, remote settings What Helps You Stand Out * Prior technical leadership on a healthcare or regulated-industry ML platform (HIPAA, HITRUST, SOC 2, or equivalent) * Direct experience with Databricks Asset Bundles, Unity Catalog, and open table formats (Iceberg, Delta) * Production experience with specialized inference pipelines beyond generic model serving (e.g., document understanding, computer vision, or streaming inference) * GPU capacity planning at strategic, tactical, and operational horizons * Experience with real-time and event-driven inference and streaming platforms (Kafka, Kinesis) * Evaluation and red-teaming experience for clinical or safety-sensitive AI * Background contributing to open-source ML infrastructure or publishing on production ML systems ## Description * Set technical direction across ML training, serving, and observability, and be the final escalation point for the most elusive infrastructure problems (GPU capacity, Spark tuning, production incidents) * Own and evolve our paved-road framework (the shared CI/CD spine, model-workflow scaffolding, and Databricks Asset Bundles) so Data Science teams can go from a config file to a production workflow without bespoke plumbing * Lead architecture for LLM-endpoint serving across managed providers (Databricks, AWS, Snowflake) and self-hosted deployments, covering latency, cost, caching, evaluation, and PHI-safe routing * Own the standards and tooling for MLflow, model registry, training image supply chain, and observability across training and inference * Partner closely with your Data & ML Platform teammates to present a cohesive ML platform to the Data Science, App Dev, and Operations teams at Datavant * Serve as a key technical input to vendor and platform selection decisions across model providers, ML tooling, and observability * Mentor senior engineers on the team, provide technical guidance to platform consumers, and stay hands-on writing high-leverage code and Infrastructure-as-Code alongside your teammates ## Related Videos - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)