> Markdown version of [/jobs/ext/2850269-staff-engineer-machine-learning-systems-reliability-moveworks](https://www.wearedevelopers.com/jobs/ext/2850269-staff-engineer-machine-learning-systems-reliability-moveworks). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Engineer, Machine Learning Systems & Reliability - Moveworks - **Company:** ServiceNow - **Location:** Mountain View, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), C++ (Programming Language), Cloud Computing, Continuous Integration, Python (Programming Language), Performance Tuning, Reliability Engineering, Site Reliability Engineering Practices, Software Engineering, Management of Software Versions, Data Processing, Large Language Models, Build Management, Kubernetes, Infrastructure Automation Frameworks, Build Tools, Machine Learning Operations, Data Pipelines, Servicenow, Golang - **Published:** September 11, 2026 - **Apply:** https://jobs.smartrecruiters.com/ServiceNow/744000148868799-staff-engineer-machine-learning-systems-reliability-moveworks ## About the Role * A track record of Staff-level technical ownership, typically gained through 7+ years of experience in software engineering, platform engineering, SRE, production engineering, or ML infrastructure. * Strong software-engineering skills in Python and at least one production systems language such as Go, Java, C++, or Rust. * Experience designing, operating, and troubleshooting distributed production systems, including failure analysis, capacity planning, and performance optimization. * Hands-on experience with cloud infrastructure, containers and Kubernetes, infrastructure as code, CI/CD, and modern observability. * Practical understanding of the ML lifecycle-including training, evaluation, model deployment, serving, monitoring, versioning, and retraining-and the ability to collaborate effectively with applied ML engineers or researchers. * Experience distinguishing service-health problems from data-quality or model-quality problems. * Familiarity with SRE practices such as SLIs/SLOs, error budgets, sustainable on-call, incident management, and blameless postmortems. * A strong automation and internal-customer mindset: you build platforms that are reliable, understandable, and pleasant for other engineers to use. * Excellent technical judgment and communication skills, especially when navigating ambiguity and coordinating across teams during production incidents. ## Description We're looking for a hands-on Staff Engineer who can move machine-learning models, agentic workflows, and self-learning approaches from promising prototypes into secure, observable, continuously deployable production systems. This role sits at the intersection of ML systems, platform engineering, and site reliability engineering. You will partner with ML, data, product, and infrastructure teams to create a paved path from experimentation to production-and take ownership of how those systems perform and evolve once deployed. What you'll do * Design and build the production path for the complete ML lifecycle: data and feature preparation, training, experiment tracking, evaluation, artifact and model management, serving, monitoring, feedback collection, and retraining. * Build continuous-delivery workflows for models, prompts, agent workflows, data dependencies, and supporting services. Establish automated quality, safety, performance, and compatibility checks. * Implement safe rollout patterns such as shadow traffic, canaries, progressive delivery, feature flags, versioned artifacts, automated rollback, and operational kill switches. * Turn self-learning approaches into controlled production feedback loops. Build systems for collecting outcomes, validating feedback, maintaining lineage, triggering model refreshes, comparing candidates, and promoting changes under explicit guardrails. * Define and operate SLIs, SLOs, alerts, and error budgets across infrastructure, data pipelines, inference services, model quality, and product behavior. * Connect model analytics and product telemetry with traditional operational signals so teams can understand whether a problem originates in infrastructure, data, model behavior, or the surrounding product. * Improve the scalability, availability, latency, and cost efficiency of distributed training, inference, and data-processing workloads. Own capacity planning and resource optimization, including GPU resources where applicable. * Participate in production ownership across the service lifecycle: architecture reviews, deployment, on-call, incident response, blameless postmortems, and systemic remediation. * Build self-service platforms and automation that reduce operational toil and shorten the time required for ML engineers and data scientists to reach production. * Apply LLMs or agentic automation to evaluation, troubleshooting, and operational workflows where they produce reliable, measurable improvements. * Establish practical standards for cloud infrastructure, Kubernetes, infrastructure as code, observability, security, and compliance. * Provide technical leadership across ML, data, product, and platform teams, mentoring engineers and influencing architecture without relying on formal authority., We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service. ## Related Videos - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Effective Machine Learning - Managing Complexity with MLOps](https://www.wearedevelopers.com/videos/185-effective-machine-learning-managing-complexity-with-mlops) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [DevOps for Machine Learning](https://www.wearedevelopers.com/videos/179-devops-for-machine-learning) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)