Machine Learning Operations Engineer

Farmers Branch
Dallas, TX, United States
25 days ago

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Computer Programming Information Engineering Distributed Systems Apache Hadoop Monitoring of Systems Job Scheduling Python (Programming Language) Performance Tuning Azure Machine Learning Software Engineering Management of Software Versions
+10 more
Feature Engineering Pandas Pyspark Low Latency Apache Kafka Spark Streaming Slurm Machine Learning Operations Stream Processing Code Restructuring

Job description

  • Optimize and maintain large-scale feature engineering pipelines using PySpark, Pandas, and PyArrow on Hadoop-based infrastructure.
  • Refactor and modularize ML codebases to enhance reusability, maintainability, and performance.
  • Collaborate with platform teams on compute capacity planning, resource allocation, and system upgrades.
  • Integrate with existing model serving frameworks to support testing, deployment, and rollback processes.
  • Monitor and troubleshoot production ML pipelines, ensuring high reliability, low latency, and cost efficiency.
  • Contribute to internal ML platforms by sharing insights, proposing improvements, and documenting best practices.
  • Build near real-time ML pipelines using Kafka and Spark Streaming.
  • Work with AWS and SageMaker MLOps ecosystem.

Requirements

  • 6+ years of experience in software engineering, data engineering, or MLOps roles.
  • Strong programming expertise in Python, with hands-on experience in Pandas, PySpark, and PyArrow.
  • Deep understanding of the Hadoop ecosystem, distributed computing, and performance tuning.
  • Experience with CI/CD pipelines and best practices in ML environments.
  • Hands-on experience with monitoring tools for ML pipeline health and performance.
  • Strong collaboration skills with experience working in cross-functional teams (platform, data science, engineering).
  • Experience contributing to or building internal MLOps frameworks/platforms.
  • Familiarity with SLURM clusters or other distributed job schedulers.
  • Exposure to Kafka, Spark Streaming, or other real-time data processing technologies.
  • Understanding of ML lifecycle management, including versioning, deployment, and drift detection.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role β€” technically off-topic, practically not.

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray Β· WWC Europe 2026

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel Β· WWC 2024

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy Β· LIVE

5:28 min

Defining MLOps and its role in production systems

Hauke Brammer Β· WWC 2023

4:43 min

Building an anti-money laundering production architecture with Python

Stefan Donsa Stefan Donsa +1 Β· LIVE

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray Β· WWC Europe 2026

Videos

See all

Related articles

See all