Machine Learning Operations Engineer
System One
Dallas, TX, United States
2 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.dallasjobsite.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours
Job source
Tech stack
Amazon Web Services
Computer Programming
Information Engineering
Distributed Systems
Apache Hadoop
Monitoring of Systems
Job Scheduling
Python (Programming Language)
Performance Tuning
Azure Machine Learning
Software Engineering
Management of Software Versions
+11 more
Feature Engineering
Pandas
Pyspark
Information Technology
Low Latency
Apache Kafka
Spark Streaming
Slurm
Machine Learning Operations
Stream Processing
Code Restructuring
Job description
- Optimize and maintain large-scale feature engineering pipelines using PySpark, Pandas, and PyArrow on Hadoop-based infrastructure.
- Refactor and modularize ML codebases to enhance reusability, maintainability, and performance.
- Collaborate with platform teams on compute capacity planning, resource allocation, and system upgrades.
- Integrate with existing model serving frameworks to support testing, deployment, and rollback processes.
- Monitor and troubleshoot production ML pipelines, ensuring high reliability, low latency, and cost efficiency.
- Contribute to internal ML platforms by sharing insights, proposing improvements, and documenting best practices.
- Build near real-time ML pipelines using Kafka and Spark Streaming.
- Work with AWS and SageMaker MLOps ecosystem., System One, and its subsidiaries including Joulé, ALTA IT Services, CM Access, TPGS, and MOUNTAIN, LTD., are leaders in delivering workforce solutions and integrated services across North America. We help clients get work done more efficiently and economically, without compromising quality. System One not only serves as a valued partner for our clients, but we offer eligible full-time employees health and welfare benefits coverage options including medical, dental, vision, spending accounts, life insurance, voluntary plans, as well as participation in a 401(k) plan.
Requirements
- 6+ years of experience in software engineering, data engineering, or MLOps roles.
- Strong programming expertise in Python, with hands-on experience in Pandas, PySpark, and PyArrow.
- Deep understanding of the Hadoop ecosystem, distributed computing, and performance tuning.
- Experience with CI/CD pipelines and best practices in ML environments.
- Hands-on experience with monitoring tools for ML pipeline health and performance.
- Strong collaboration skills with experience working in cross-functional teams (platform, data science, engineering).
- Experience contributing to or building internal MLOps frameworks/platforms.
- Familiarity with SLURM clusters or other distributed job schedulers.
- Exposure to Kafka, Spark Streaming, or other real-time data processing technologies.
- Understanding of ML lifecycle management, including versioning, deployment, and drift detection.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dallasjobsite.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
BB
Benedikt Bischof
about 4 years ago
BB
Benedikt Bischof
MLOps – What’s the deal behind it?
almost 4 years ago
LM
Luis Minvielle
How to Become an AI Engineer
almost 3 years ago
LM
Luis Minvielle
What Are Large Language Models?
almost 3 years ago
BB
Benedikt Bischof
MLOps And AI Driven Development
over 4 years ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
23 days ago