machine learning infrastructure engineer

Motion Recruitment Partners LLC.
Woburn, MA, United States
1 day ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Computer Vision Cloud Computing Cloud Database Cloud Engineering Continuous Integration Data Files Software Debugging DevOps Github Mobile Application Software Python (Programming Language)
+26 more
Machine Learning Regression Testing Tensorflow Azure Machine Learning Software Deployment Management of Software Versions Image Acquisition Edge Inference TorchServe Google Cloud Pytorch Autoscaling On-device Inference Model Validation Discretization Infrastructure Automation Frameworks Low Latency ONNX (Open Neural Network Exchange) Format Free and Open-Source Software Machine Learning Operations TensorRT Multiaccess Edge Computing Terraform Data Pipelines Human in the Loop Docker

Job description

We are seeking a Staff MLOps Engineer to join an innovative Series B robotics and automotive technology startup in Woburn, MA. This is a full-time opportunity for a senior machine learning infrastructure engineer who wants to own the production lifecycle of computer vision models powering real-world automotive service applications. You’ll work across Google Cloud Platform, Docker, Python, PyTorch/TensorFlow, model serving, CI/CD, infrastructure as code, and scalable inference systems.

This is a true ownership role where you’ll take an existing machine learning product from MVP to production scale. You’ll own the architecture behind how computer vision models are deployed, monitored, evaluated, improved, and served to customers. The technology is already being used in real automotive service environments, so your decisions will have a direct impact on product performance, customer experience, and the company’s ability to scale. You’ll also serve as a senior cloud architecture voice alongside the broader engineering team, making this a great opportunity for someone who wants significant technical ownership without stepping away from hands-on engineering., Own the architecture and operation of the company’s multi-stage computer vision inference pipeline Re-architect the current MVP infrastructure into a scalable production serving platform Design containerized inference infrastructure with GPU acceleration, queueing, batching, and autoscaling where appropriate Own model deployment, versioning, staged rollouts, canary testing, shadow evaluation, and rollback Build the data and model improvement lifecycle from field data collection through labeling, training, evaluation, and deployment Establish dataset versioning and reproducible evaluation processes Monitor model performance and identify drift or degradation in production Analyze segmented model performance and investigate real-world failures Define and monitor commercial ML metrics including false positives, false negatives, technician overrides, latency, and inference cost Improve model performance through architecture selection, augmentation, hard-example mining, quantization, and distillation Evaluate cloud versus on-device inference and help determine the right architecture Build MLOps foundations including experiment tracking, reproducible training, model CI/CD, and infrastructure as code Partner with senior software engineers to establish cloud architecture and Google Cloud Platform best practices Work with hardware and field teams to improve image capture quality, including lighting, focus, and probe positioning Own production support and incident response for the ML platform Participate in on-call responsibilities for inference availability Identify technical risks and communicate architectural trade-offs to engineering and company leadership

Tech Breakdown

25% MLOps / Model Deployment & Serving 20% Google Cloud Platform / Cloud Infrastructure 20% Computer Vision / ML Engineering 15% Data Pipelines / Evaluation / Model Improvement 10% DevOps / Infrastructure as Code / CI/CD 10% Architecture / Technical Leadership

Daily Responsibilities

45% Hands-On Engineering 20% ML Infrastructure & Production Operations 15% Model Performance & Evaluation 10% Architecture & Technical Leadership 10% Cross-Functional Collaboration

Requirements

8+ years of professional engineering experience Several years of experience owning machine learning systems in production Proven experience taking computer vision or ML pipelines from prototype through production scale Deep experience with segmentation and classification models, preferably in multi-stage pipelines Strong cloud infrastructure experience, preferably with Google Cloud Platform Experience with Vertex AI, Cloud Run, GKE, Cloud Functions, Cloud SQL, Pub/Sub, and Docker Strong Python development skills Production experience with PyTorch or TensorFlow Hands-on experience with model serving and optimization Experience with technologies such as Triton, TorchServe, ONNX, TensorRT, quantization, or similar platforms Experience owning production deployments, incident response, rollback procedures, and on-call support Experience building data collection, labeling, and dataset versioning workflows Experience developing reproducible model evaluation and regression testing systems Experience with Terraform or similar infrastructure-as-code tools Experience with CI/CD automation such as GitHub Actions Strong understanding of the trade-offs between model accuracy, latency, infrastructure cost, and reliability Excellent debugging, problem-solving, and communication skills

Desired Skills & Experience

Experience with on-device or edge inference Experience with Core ML, TensorFlow Lite, or ExecuTorch Experience building active learning or human-in-the-loop labeling systems Computer vision experience with smaller, long-tail, or industrial inspection datasets Experience integrating cameras, sensors, or other hardware with ML systems Robotics, IoT, or edge computing experience Experience with ROS or similar robotics platforms Experience integrating ML capabilities into mobile applications Automotive service, dealership, DMS, or automotive technology experience Experience working in an early-stage or high-growth startup Open-source contributions Experience collaborating closely with hardware and field operations teams

Benefits & conditions

Competitive salary Comprehensive benefits package Opportunity to own the ML platform for a product already being used by paying customers Significant technical ownership over the company’s production architecture Opportunity to work with real-world proprietary computer vision data Hands-on exposure to robotics, automotive technology, computer vision, and edge computing Work alongside experienced robotics and software engineering professionals Collaborative, low-ego, high-intensity startup environment Prime Woburn, MA location with on-site parking Opportunity to have a direct impact on the company’s ability to scale

You will receive the following benefits:

Medical Insurance Dental Benefits Vision Benefits Paid Time Off (PTO) 401(k) Comprehensive Benefits Package Professional Development Opportunities Collaborative Startup Environment On-Site Parking

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Loading talks and stories from around this role…