Machine Learning Operations Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+11 more
Job description
We are seeking a Machine Learning Operations (MLOps) Engineer to join our team. The MLOps Engineer will be responsible for building and maintaining the infrastructure that enables reliable deployment, monitoring, governance, and continuous improvement of production machine learning systems across enterprise client environments. What You’ll Do:
- Design, build, and maintain scalable machine learning deployment pipelines.
- Develop standardized model registries, artifact repositories, data versioning, and reproducible ML environments.
- Build automated evaluation pipelines for production machine learning models.
- Implement automated data quality monitoring including profiling, anomaly detection, validation, quarantine, and alerting.
- Develop automated retraining workflows, promotion gates, rollback capabilities, and audit trails.
- Monitor production environments for model drift, latency, prediction quality, infrastructure performance, and operational costs.
- Troubleshoot production machine learning issues and lead incident response activities.
- Build CI/CD pipelines supporting enterprise AI applications.
- Collaborate closely with Data Scientists and client engineering teams to deploy and maintain AI solutions.
- Ensure governance, security, lineage, reproducibility, and audit readiness across machine learning platforms.
Requirements
- 8+ years - Passionate about building reliable AI infrastructure at enterprise scale.
- Experienced deploying and maintaining production machine learning systems.
- Strong analytical and troubleshooting skills.
- Fast learner with attention to detail.
- Excellent communication and collaboration skills.
- Comfortable working with both software engineering and data science teams.
Education: Bachelor’s or Master’s degree in Computer Science, Software Engineering, Data Engineering, or a related technical field. Related Work Experience: 3+ years supporting production machine learning platforms or cloud infrastructure. Technical Skills:
- Advanced SQL
- Python
- Apache Spark / PySpark
- AWS SageMaker (Databricks, Azure ML, or Vertex AI experience is a plus)
- Kubernetes
- CI/CD pipelines
- Infrastructure as Code (Terraform, CloudFormation, or similar)
- Model monitoring and ML observability tools
- Data validation and automated testing frameworks
-
Statistics related to monitoring, model drift, and performance evaluation
-
- Git and modern DevOps practices
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps And AI Driven Development
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production
How to Become an AI Engineer
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production