> Markdown version of [/jobs/ext/2691989-machine-learning-operations-engineer](https://www.wearedevelopers.com/jobs/ext/2691989-machine-learning-operations-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Operations Engineer - **Company:** System One - **Location:** Dallas, TX, United States - **Experience:** Expert - **Contract:** Temporary to permanent - **Skills:** Amazon Web Services, Computer Programming, Information Engineering, Distributed Systems, Apache Hadoop, Monitoring of Systems, Job Scheduling, Python (Programming Language), Performance Tuning, Azure Machine Learning, Software Engineering, Management of Software Versions, Feature Engineering, Pandas, Pyspark, Information Technology, Low Latency, Apache Kafka, Spark Streaming, Slurm, Machine Learning Operations, Stream Processing, Code Restructuring - **Published:** September 3, 2026 - **Apply:** https://www.dallasjobsite.com/job.asp?id=3375502569&tx=JK6559FFI&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role * 6+ years of experience in software engineering, data engineering, or MLOps roles. * Strong programming expertise in Python, with hands-on experience in Pandas, PySpark, and PyArrow. * Deep understanding of the Hadoop ecosystem, distributed computing, and performance tuning. * Experience with CI/CD pipelines and best practices in ML environments. * Hands-on experience with monitoring tools for ML pipeline health and performance. * Strong collaboration skills with experience working in cross-functional teams (platform, data science, engineering). * Experience contributing to or building internal MLOps frameworks/platforms. * Familiarity with SLURM clusters or other distributed job schedulers. * Exposure to Kafka, Spark Streaming, or other real-time data processing technologies. * Understanding of ML lifecycle management, including versioning, deployment, and drift detection. ## Description * Optimize and maintain large-scale feature engineering pipelines using PySpark, Pandas, and PyArrow on Hadoop-based infrastructure. * Refactor and modularize ML codebases to enhance reusability, maintainability, and performance. * Collaborate with platform teams on compute capacity planning, resource allocation, and system upgrades. * Integrate with existing model serving frameworks to support testing, deployment, and rollback processes. * Monitor and troubleshoot production ML pipelines, ensuring high reliability, low latency, and cost efficiency. * Contribute to internal ML platforms by sharing insights, proposing improvements, and documenting best practices. * Build near real-time ML pipelines using Kafka and Spark Streaming. * Work with AWS and SageMaker MLOps ecosystem., System One, and its subsidiaries including Joulé, ALTA IT Services, CM Access, TPGS, and MOUNTAIN, LTD., are leaders in delivering workforce solutions and integrated services across North America. We help clients get work done more efficiently and economically, without compromising quality. System One not only serves as a valued partner for our clients, but we offer eligible full-time employees health and welfare benefits coverage options including medical, dental, vision, spending accounts, life insurance, voluntary plans, as well as participation in a 401(k) plan. ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Effective Machine Learning - Managing Complexity with MLOps](https://www.wearedevelopers.com/videos/185-effective-machine-learning-managing-complexity-with-mlops) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)