Sr. Machine Learning Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+17 more
Job description
This position owns the production platform and operationalization of machine learning models developed by Data Scientists. The Senior MLOps Engineer designs, develops, deploys, secures, monitors, scales, and maintains the AWS- and Kubernetes-based infrastructure, CI/CD pipelines, and supporting software services required to run approved models reliably in production. The role covers existing products and new products in development. Data Scientists retain ownership of model development, experimentation, training, evaluation, and selection; the Senior MLOps Engineer enables those models to be deployed, integrated, operated, benchmarked, and improved safely at scale.
What You’ll Do
- Lead complex MLOps projects - Drive highly complex engineering initiatives with broad autonomy and independent judgment.
- Design & develop software - Build, document, test, and debug software solutions for both customer-facing and internal applications.
- Operate production ML services - Design, deploy, scale, secure, monitor, troubleshoot, and maintain production systems hosting Data Scientist-developed models.
- Kubernetes & AWS architecture - Build and maintain Kubernetes deployments, AWS infrastructure, container images, networking, secrets, IAM, and observability for model-serving workloads.
- CI/CD & IaC workflows - Develop and maintain CI/CD pipelines and infrastructure-as-code workflows for model-serving systems and supporting components.
- API & orchestration development - Create, test, document, and maintain APIs, orchestration layers, integrations, and data/service interfaces required for production model operations.
- Partner with Data Scientists - Productionize approved models, define deployment requirements, and resolve integration issues; model development and training remain with Data Science.
- Benchmarking & performance testing - Own benchmarking, capacity planning, load testing, performance testing, and reliability testing for ML, LLM, and generative-model services.
- Operational standards - Establish and maintain standards for logging, metrics, alerting, dashboards, incident response, rollback, availability, security, and cost-efficient operation.
- MLflow & Kubeflow operations - Operate and enhance MLflow and Kubeflow capabilities for model lifecycle, pipelines, deployment, and operational workflows.
- Evaluate new technologies - Identify and assess new platforms and technologies to strengthen existing systems.
- Enforce engineering standards - Follow, define, and enforce development best practices and engineering standards.
- Mentor team members - Provide guidance and support to less experienced engineers.
- Collaborate on roadmaps - Work with architects and engineering managers to maintain development roadmaps and prioritize features.
- Cross-team engagement - Engage with data science and engineering teams to understand requirements and constraints across multiple products.
- Share MLOps expertise - Actively share technical knowledge and MLOps framework expertise with teammates.
- Stay current in MLOps - Maintain up-to-date knowledge of MLOps and data-science technologies.
- Cybersecurity domain expertise - Develop domain expertise in at least one cybersecurity application area.
- Write technical documentation
Requirements
- 7+ years engineering experience - Software engineering, platform engineering, DevOps, SRE, MLOps, or related roles, including 1+ year in a senior position and substantial hands-on production operations work.
- Deep software & cloud knowledge - Strong understanding of software engineering, cloud-platform engineering, DevOps, and MLOps practices; experience leading projects that deliver and operate production systems.
- Production architecture & API design - Extensive experience designing production architectures and building reliable APIs, integrations, automation, and supporting services in Python; Java or C++ is a plus.
- Comprehensive MLOps background - Containerized model serving, model lifecycle tooling, CI/CD, IaC, observability, release management, and secure operation of production services. (Model development/training stays with Data Science.)
- AWS & Kubernetes operations - Hands-on deployment and operation of applications using Docker, IAM, secrets management, networking, monitoring, logging, alerting, and incident troubleshooting.
- CI/CD pipeline ownership - Required experience building and maintaining CI/CD pipelines; Jenkins and ArgoCD are strong pluses.
- ML model service deployment - Proven experience deploying, operating, benchmarking, and load-testing machine-learning model services in production.
- MLflow & Kubeflow experience - Strong preference for hands-on experience; candidates without prior exposure must be able to learn and apply both quickly.
- ML framework familiarity - Knowledge of PyTorch, TensorFlow, scikit-learn, etc., for integrating and serving models built by Data Scientists (not responsible for model creation).
- Cross-functional collaboration - Demonstrated ability to work with engineering leads, software architects, and other stakeholders to advance engineering and data-science initiatives.
- Project leadership - Successfully led multiple projects to completion within small teams.
- Mentorship experience - Proven ability to assist or mentor junior engineers.
- Strong communication skills - Able to clearly explain complex technical topics to non-technical audiences, verbally and in writing.
- Problem-solving & critical thinking - Demonstrated excellence in tackling diverse engineering challenges.
- Knowledge sharing - Comfortable and enthusiastic about sharing technical expertise with teammates.
- Cybersecurity interest
About the company
Whether you’re an experienced professional or just getting started, your contributions matter at Fortra. If you’re passionate about tackling meaningful challenges alongside talented team members committed to helping each other succeed, all while having lots of fun, we want to hear from you. We offer competitive benefits and salaries, personal and professional development opportunities, flexibility, and much more!
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path
Why Upskilling And Reskilling is Important For Developers