Machine Learning Ops Engineer

Horizontal Talent
Boston, MA, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours

Tech stack

Amazon Web Services Bash Shell Computer Programming Continuous Integration Information Engineering Data Infrastructure DevOps Python (Programming Language) Machine Learning Cloud Services Software Engineering Large Language Models
+7 more
Containerization Infrastructure Automation Frameworks Information Technology Machine Learning Operations Stream Processing Api Management Docker

Job description

We are seeking a dedicated and innovative Machine Learning Ops Engineer to join our dynamic team. This role is essential for bridging the gap between machine learning, software engineering, and platform operations, contributing to impactful solutions in a collaborative environment. Responsibilities

  • Design and implement automated ML pipelines for model training, evaluation, and deployment.
  • Build and manage scalable model serving infrastructure for real-time and batch scoring.
  • Architect and operate a sub-model orchestration layer compliant with regulatory requirements.
  • Design and maintain a feature store architecture to ensure training-serving consistency.
  • Implement and govern LLM API integrations, ensuring effective prompt management and cost tracking.
  • Establish production monitoring and alerting systems for model performance and pipeline integrity.
  • Automate infrastructure provisioning using Infrastructure-as-Code practices.
  • Define and lead the enterprise MLOps platform strategy, mentoring junior engineers along the way.
  • Establish CI/CD workflows to facilitate rapid and safe ML iteration.
  • Engage in other duties and special projects as assigned.

Requirements

  • Bachelor’s or Master’s Degree in Computer Science, Software Engineering, or a related technical field.
  • 7+ years of professional experience in DevOps, Data Engineering, or ML Engineering, with a focus on Machine Learning operations.
  • Expertise in containerization and orchestration technologies such as Docker and Kubernetes.
  • Proficiency in ML lifecycle tooling and cloud services, particularly AWS.
  • Strong programming skills in Python, with experience in Bash scripting.

Preferred Skills

  • Experience in healthcare or regulated data environments, with familiarity in HIPAA compliance.
  • Knowledge of stream processing and event-driven pipeline design.
  • Experience with LLM API integration and management.
  • Strong interpersonal skills and ability to communicate technical concepts clearly.
  • A proactive and curious mindset, eager to adopt emerging MLOps tools and practices.

We believe in fostering a diverse, equitable, and inclusive workplace where all team members feel valued and empowered to contribute their unique perspectives. We encourage applicants from all backgrounds to apply and join us in our mission to create innovative solutions.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on disabledperson.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:28 min

Defining MLOps and its role in production systems

Hauke Brammer · WWC 2023

3:52 min

Avoiding remote code execution from unsanitized inputs

Alexander Pirker · WWC 2022

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

4:19 min

Introduction to DevOps for AI and MLOps

Aarno Aukia · LIVE

54 sec

Interpreting complex terminal commands safely using external explanation utilities

Dan Cranney +2 · LIVE

Videos

See all

Related articles

See all