> Markdown version of [/jobs/ext/3606757-machine-learning-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/3606757-machine-learning-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # machine learning infrastructure engineer - **Company:** Motion Recruitment Partners LLC. - **Location:** Woburn, MA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computer Vision, Cloud Computing, Cloud Database, Cloud Engineering, Continuous Integration, Data Files, Software Debugging, DevOps, Github, Mobile Application Software, Python (Programming Language), Machine Learning, Regression Testing, Tensorflow, Azure Machine Learning, Software Deployment, Management of Software Versions, Image Acquisition, Edge Inference, TorchServe, Google Cloud, Pytorch, Autoscaling, On-device Inference, Model Validation, Discretization, Infrastructure Automation Frameworks, Low Latency, ONNX (Open Neural Network Exchange) Format, Free and Open-Source Software, Machine Learning Operations, TensorRT, Multiaccess Edge Computing, Terraform, Data Pipelines, Human in the Loop, Docker - **Published:** October 7, 2026 - **Apply:** https://www.dice.com/job-detail/7faf7bb8-56c1-4684-b6f7-dc4febef2537 ## About the Role 8+ years of professional engineering experience Several years of experience owning machine learning systems in production Proven experience taking computer vision or ML pipelines from prototype through production scale Deep experience with segmentation and classification models, preferably in multi-stage pipelines Strong cloud infrastructure experience, preferably with Google Cloud Platform Experience with Vertex AI, Cloud Run, GKE, Cloud Functions, Cloud SQL, Pub/Sub, and Docker Strong Python development skills Production experience with PyTorch or TensorFlow Hands-on experience with model serving and optimization Experience with technologies such as Triton, TorchServe, ONNX, TensorRT, quantization, or similar platforms Experience owning production deployments, incident response, rollback procedures, and on-call support Experience building data collection, labeling, and dataset versioning workflows Experience developing reproducible model evaluation and regression testing systems Experience with Terraform or similar infrastructure-as-code tools Experience with CI/CD automation such as GitHub Actions Strong understanding of the trade-offs between model accuracy, latency, infrastructure cost, and reliability Excellent debugging, problem-solving, and communication skills Desired Skills & Experience Experience with on-device or edge inference Experience with Core ML, TensorFlow Lite, or ExecuTorch Experience building active learning or human-in-the-loop labeling systems Computer vision experience with smaller, long-tail, or industrial inspection datasets Experience integrating cameras, sensors, or other hardware with ML systems Robotics, IoT, or edge computing experience Experience with ROS or similar robotics platforms Experience integrating ML capabilities into mobile applications Automotive service, dealership, DMS, or automotive technology experience Experience working in an early-stage or high-growth startup Open-source contributions Experience collaborating closely with hardware and field operations teams ## Description We are seeking a Staff MLOps Engineer to join an innovative Series B robotics and automotive technology startup in Woburn, MA. This is a full-time opportunity for a senior machine learning infrastructure engineer who wants to own the production lifecycle of computer vision models powering real-world automotive service applications. You'll work across Google Cloud Platform, Docker, Python, PyTorch/TensorFlow, model serving, CI/CD, infrastructure as code, and scalable inference systems. This is a true ownership role where you'll take an existing machine learning product from MVP to production scale. You'll own the architecture behind how computer vision models are deployed, monitored, evaluated, improved, and served to customers. The technology is already being used in real automotive service environments, so your decisions will have a direct impact on product performance, customer experience, and the company's ability to scale. You'll also serve as a senior cloud architecture voice alongside the broader engineering team, making this a great opportunity for someone who wants significant technical ownership without stepping away from hands-on engineering., Own the architecture and operation of the company's multi-stage computer vision inference pipeline Re-architect the current MVP infrastructure into a scalable production serving platform Design containerized inference infrastructure with GPU acceleration, queueing, batching, and autoscaling where appropriate Own model deployment, versioning, staged rollouts, canary testing, shadow evaluation, and rollback Build the data and model improvement lifecycle from field data collection through labeling, training, evaluation, and deployment Establish dataset versioning and reproducible evaluation processes Monitor model performance and identify drift or degradation in production Analyze segmented model performance and investigate real-world failures Define and monitor commercial ML metrics including false positives, false negatives, technician overrides, latency, and inference cost Improve model performance through architecture selection, augmentation, hard-example mining, quantization, and distillation Evaluate cloud versus on-device inference and help determine the right architecture Build MLOps foundations including experiment tracking, reproducible training, model CI/CD, and infrastructure as code Partner with senior software engineers to establish cloud architecture and Google Cloud Platform best practices Work with hardware and field teams to improve image capture quality, including lighting, focus, and probe positioning Own production support and incident response for the ML platform Participate in on-call responsibilities for inference availability Identify technical risks and communicate architectural trade-offs to engineering and company leadership Tech Breakdown 25% MLOps / Model Deployment & Serving 20% Google Cloud Platform / Cloud Infrastructure 20% Computer Vision / ML Engineering 15% Data Pipelines / Evaluation / Model Improvement 10% DevOps / Infrastructure as Code / CI/CD 10% Architecture / Technical Leadership Daily Responsibilities 45% Hands-On Engineering 20% ML Infrastructure & Production Operations 15% Model Performance & Evaluation 10% Architecture & Technical Leadership 10% Cross-Functional Collaboration