> Markdown version of [/jobs/ext/2680256-machine-learning-operations-engineer](https://www.wearedevelopers.com/jobs/ext/2680256-machine-learning-operations-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Operations Engineer - **Company:** The University of Texas MD Anderson Cancer Center - **Location:** Houston, TX, United States (Remote available) - **Experience:** Expert - **Salary:** $183,000.0 - $219,500.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Computer Vision, Microsoft Azure, Health Informatics, Cloud Computing, Directed Acyclic Graph (Directed Graphs), Information Engineering, DevOps, Github, Machine Learning, NumPy, Azure Machine Learning, Software Engineering, Enterprise Software Applications, Pytorch, Containerization, Kubernetes, Information Technology, Machine Learning Operations, Software Version Control, Data Pipelines, Docker - **Published:** September 2, 2026 - **Apply:** https://jobs.mdanderson.org/about/apply ## About the Role The ideal candidate is a seasoned machine learning or software engineering professional with a strong foundation in MLOps, cloud and on-premises AI platforms, and healthcare-focused AI lifecycle management. This individual typically holds a Bachelor's degree in a relevant technical discipline, with a Master's degree preferred, and brings significant hands-on experience developing, deploying, and maintaining machine learning systems in production environments. Experience leading or designing shared ML services, evaluating third-party AI solutions, and applying responsible AI practices within regulated or clinical settings is highly valued., Education Required: Bachelor's degree in Computer Science, Software Engineering, Data Science, Physics, Math & Statistics, or another related engineering discipline. Preferred Education: Master's Level Degree Experience Required : Five years of experience in machine learning engineering, data science, data engineering, and/or software engineering. With Master's degree, three years' experience required. With PhD, one year of experience required. Preferred Experience: Experience developing MLOps pipelines for computer vision AI models, hands on experience developing custom machine learning algorithms from scratch (e.g., in NumPy or PyTorch, designed and implemented shared machine learning service that is used across multiple teams or production projects, led the development of systems that automate the deployment and maintenance of multiple machine learning models into user-facing products, five years of industry experience in data science, with at least 3 of those years as a Senior Machine Learning Engineer ## Description Within this mission-driven environment, the Senior Machine Learning Operations Engineer plays a critical role in building, deploying, and sustaining production-quality machine learning systems. The Senior Machine Learning Operations Engineer partners closely with data scientists, engineers, clinicians, and business stakeholders to ensure AI solutions are scalable, secure, reliable, and aligned with responsible AI principles across UT MD Anderson., AI Model Lifecycle & MLOps * Oversee end-to-end AI model lifecycles including training, evaluation, deployment, monitoring, and maintenance of production-quality machine learning models * Design and implement CI/CD pipelines for model training, deployment, monitoring, and retraining with a focus on security, scalability, reliability, reproducibility, and performance * Implement rigorous testing, versioning, and documentation practices to support reproducibility, risk mitigation, and measurable impact * Maintain comprehensive experiment tracking, data lineage, model lineage, and model scorecards * Design fallback, rollback, and decommissioning strategies to ensure operational continuity of AI solutions Responsible AI & Governance * Promote responsible AI practices by minimizing bias, enhancing fairness, and maximizing transparency in machine learning models * Ensure AI lifecycle management aligns with institutional standards and best practices * Support assessment, validation, and onboarding of external machine learning models and AI-driven products to minimize organizational risk and maximize value Platform, Infrastructure & Tooling * Develop and maintain scalable data pipelines, feature stores, and artifact management systems * Deploy and operate ML workloads across cloud and on-premises environments including Azure, AWS, or GCP * Utilize containerization and orchestration technologies such as Docker, Kubernetes, and DAG-based tools * Apply DevOps and MLOps tools including Azure DevOps, GitHub Actions, and version control systems Stakeholder Engagement & Enablement * Collaborate with stakeholders to gather requirements, translate AI concepts into understandable terms, and incorporate feedback * Partner with data scientists, ML engineers, and software engineers to integrate models into enterprise systems * Deliver training and knowledge sharing to enhance AI understanding and adoption across the organization * Report project progress, impact, risks, and recommendations to leadership Innovation & Continuous Learning * Stay current with emerging technology trends in AI, MLOps, and healthcare analytics * Contribute to internal and external technical communities * Foster a culture of continuous improvement, innovation, and learning across teams * Perform other duties as assigned, This position may be responsible for maintaining the security and integrity of critical infrastructure, as defined in Section 113.001(2) of the Texas Business and Commerce Code and therefore may require routine reviews and screening. The ability to satisfy and maintain all requirements necessary to ensure the continued security and integrity of such infrastructure is a condition of hire and continued employment. ## Related Videos - [Effective Machine Learning - Managing Complexity with MLOps](https://www.wearedevelopers.com/videos/185-effective-machine-learning-managing-complexity-with-mlops) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [DevOps for Machine Learning](https://www.wearedevelopers.com/videos/179-devops-for-machine-learning) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)