Machine Learning Ops (MLOps) Engineer to architect

The Mosaic Company
Nashville, TN, United States
6 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Business Analytics Applications Microsoft Azure Batch Processing Cloud Computing Cloud Database Continuous Integration Information Engineering Python (Programming Language) Machine Learning Regression Testing Datadog
+13 more
Data Logging System Availability Large Language Models Snowflake Technical Debt Jupyter Containerization Information Technology Machine Learning Operations Code Restructuring Data Pipelines Workday Docker

Job description

We are seeking an experienced Machine Learning Ops (MLOps) Engineer to architect, develop, and maintain the full lifecycle of data and model pipelines that power training, inference, evaluation, and analytics workflows. This role is responsible for ensuring the reliability, scalability, and observability of all machine learning systems in production, including traditional ML models and modern LLM-based/MCP-orchestrated architectures. A key focus of this role in the near term is auditing and consolidating our existing pipelines and deployment processes. The ideal candidate is highly skilled in Python, Jupyter, Snowflake, and both Azure and AWS cloud environments, and thrives in environments requiring continuous monitoring, rapid issue diagnosis, and rigorous validation before deployment., * Design, build, and maintain scalable data pipelines supporting model training, inference, batch processing, and real-time analytics workflows.

  • Audit, refactor, and consolidate existing ML pipelines and deployment processes to eliminate technical debt, redundant workflows, and undocumented manual steps.
  • Audit, refactor, and consolidate existing ML pipelines and deployment processes to eliminate technical debt, redundant workflows, and undocumented manual steps.
  • Monitor and deploy and deploy production ML pipelines to identify anomalies, performance degradations, or failures related to data quality, logic defects, or infrastructure issues.
  • Execute rapid troubleshooting and root-cause analysis followed by timely remediation, validation, and full regression testing prior to redeployment.
  • Collaborate with Data Science, Engineering, and Product teams to operationalize machine learning models-including LLM-based and MCP-orchestrated systems-ensuring seamless integration into production environments.
  • Develop CI/CD workflows, model deployment strategies, and automated testing frameworks to support reliable, repeatable releases.
  • Implement and maintain observability tooling (logging, monitoring, alerting) to ensure high availability and traceability of ML systems.
  • Manage and optimize cloud infrastructure across Azure and AWS for compute, storage, orchestration, and security needs.
  • Create and maintain documentation, runbooks, and best practices for model operations and system maintenance.
  • Perform all other job-related duties as assigned., * This position uses a computer and other office equipment as needed to perform duties. The in-office noise level in the work environment is typical of that of an office. Frequent interruptions may be encountered throughout the workday.
  • The employee is required to either stand or sit, talk and hear frequently required to use repetitive keying or hand motions.
  • The physical demands are representative of those that must be met by an employee to successfully perform the essential functions of this job. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.

Requirements

  • Bachelor’s Degree in Computer Science, Engineering or equivalent work experience.
  • 5-7 years of combined experience in Data Engineering, MLOps, Machine Learning Engineering, or related fields.
  • Demonstrated experience operationalizing traditional ML models as well as LLM-based and MCP-orchestrated systems.
  • Strong working knowledge of both Azure and AWS cloud platforms, including compute orchestration, networking, and security best practices.
  • Experience with CI/CD tools, containerization (Docker), infrastructure-as-code, and ML pipeline frameworks.
  • Strong ability to diagnose and resolve pipeline failures, data anomalies, and complex system issues.

Advanced proficiency in Python, Jupyter, and common ML/analytics frameworks.

  • Hands-on experience with Snowflake or similar cloud data warehousing environment.
  • Excellent problem-solving skills, attention to detail, and a proactive, self-directed work ethic.
  • Strong communication skills and comfort working in fast-paced, cross-functional environments.

Work Environment

  • This role is preferred to be based in Nashville or Jacksonville, near Mosai’s offices.

About the company

Mosai is the intelligent care coordination platform that brings together the fragmented pieces of healthcare into a clear, connected picture. Like a mosaic, our platform unites data, people, and processes so providers can make better decisions, coordinate care in real time, and deliver improved outcomes. With Mosai, home-based care organizations can thrive in value-based care while giving every patient the right care, in the right place, at the right time. Learn more at ;br>

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:28 min

Defining MLOps and its role in production systems

Hauke Brammer · World Congress 2023

3:31 min

Producing candlestick visualization charts inside integrated Jupyter notebooks

Akmal Chaudhri Akmal Chaudhri · LIVE

6:22 min

Eliminating HR bureaucracy and trusting employees

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

5:36 min

Building a data analysis stack with Python and Jupyter

Markus Harrer Markus Harrer · World Congress 2021

5:06 min

Primary reasons for capability gaps in modern recruitment systems

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

Videos

See all

Related articles

See all