Machine Learning Engineer (Remote)

RECRUITER LLC
United States
1 day ago
Apply on arc.dev
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$175,000.0 - $225,000.0
Working hours
Regular working hours
Job source

Tech stack

A/B Testing Airflow Computer Vision Continuous Integration Github Python (Programming Language) Machine Learning Software Engineering Delivery Pipeline Reliability of Systems Apache Flink Deployment Automation
+5 more
Machine Learning Operations Api Design Terraform Docker Microservices

Job description

This opportunity is for a Senior Machine Learning Engineer, Forecasting focused on production machine learning systems, MLOps infrastructure, and forecasting workflows.

The position owns critical parts of the production ML lifecycle, including training automation, deployment pipelines, model serving infrastructure, monitoring, and scalable systems that enable reliable model development and iteration.

The role combines Python software engineering, distributed ML systems, Docker and Kubernetes, CI/CD, infrastructure-as-code, model orchestration, and automated retraining. It works closely with engineering, product, and data science teams to productionize machine learning research, improve experimentation velocity, investigate model and data-quality issues, and strengthen the reliability and maintainability of production forecasting systems.

What You’ll Do

  • Collaborate with engineering, product, and data science teams to understand business challenges and identify opportunities for machine learning and AI solutions.
  • Develop tools and automate manual processes to improve operational efficiency, accelerate experimentation, and reduce human error.
  • Build, integrate, and monitor end-to-end lifecycles for large-scale, distributed machine learning systems.
  • Investigate model performance and identify data-quality and system-performance issues.
  • Improve the ML pipeline supporting the forecasting platform, including weekly automated model retraining and deployment across multiple production models.
  • Advance MLOps best practices, deployment automation, and production ML capabilities across the team.

Requirements

Experience: 5+ years building and maintaining production ML systems

Core Areas: MLOps, production ML systems, forecasting, model serving infrastructure, automated retraining, ML deployment automation, distributed machine learning systems, model monitoring

Compensation: $175,000 - $225,000 per year, plus equity, * 5+ years of experience building and maintaining production ML systems, with deep expertise in MLOps, deployment automation, and model serving infrastructure.

  • Strong software engineering skills and proficiency in Python.
  • Hands-on experience with containerization and orchestration technologies, including Docker and Kubernetes.
  • Experience with CI/CD systems such as GitHub Actions and ArgoCD.
  • Experience with infrastructure-as-code and deployment tooling such as Terraform and Helm.
  • Production ML deployment experience, including model training orchestration with Dagster, Airflow, or similar tools.
  • Experience building automated model retraining pipelines and working with A/B testing and variant management.
  • Systems design expertise, including scalable microservices, API design, and management of complex service dependencies.
  • Demonstrated ownership of projects from concept through production, including ongoing maintenance and continuous improvements to system reliability.
  • Ability to work effectively with data scientists to productionize research, backend engineering teams on API integration, and product teams to address customer needs.
  • Commitment to writing tested, maintainable, and well-documented code that supports team velocity.

Preferred Qualifications

  • Experience in healthcare or another regulated industry.
  • Experience with forecasting or time series algorithms.
  • Experience with computer vision.
  • Experience with DAG frameworks or Flink.

Interview Process

  • Initial conversation with a recruiter to discuss background and the role.
  • Interview with the hiring manager to review relevant experience and complete a role-related case or technical exercise.
  • Virtual onsite interviews with several team members covering collaboration, culture, machine learning engineering coding skills, and system design. This stage typically consists of three to five interviews.

Benefits & conditions

  • Stock options.
  • Flexible vacation policy.
  • Remote-first work environment with virtual and in-person events designed to support team connection.
  • Comprehensive health, dental, and vision insurance.
  • 16 weeks of parental leave for all parents.
  • Professional growth opportunities.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on arc.dev
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

1:30 min

Challenges of automated application screening in recruiting

Kilian Kluge +1 · World Congress 2022

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

Videos

See all

Related articles

See all