> Markdown version of [/jobs/ext/2278910-senior-machine-learning-engineer-agent-eval-platform](https://www.wearedevelopers.com/jobs/ext/2278910-senior-machine-learning-engineer-agent-eval-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Machine Learning Engineer, Agent Eval Platform - **Company:** ServiceNow - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Python (Programming Language), Machine Learning, Large Language Models, Prompt Engineering, Servicenow - **Published:** August 28, 2026 - **Apply:** https://jobs.smartrecruiters.com/ServiceNow/744000146022698-senior-machine-learning-engineer-agent-eval-platform ## About the Role To be successful in this role you have: * 5+ years in applied ML, data science, or ML-adjacent engineering, with a track record of work that shipped and got used * Experience turning subjective human judgement into a measurement that holds up - one that other people, and ideally other models, can act on. This is the core of the job * Strong applied ML fundamentals, and comfort treating LLMs as a component you evaluate, prompt, and fine-tune rather than one you pretrain * Strong Python, and the discipline to ship production-grade code rather than notebooks * Ability to think and communicate clearly about complex problems - a large part of this job is convincing engineers that a number means what you say it means, and being right * A high degree of ownership and a bias toward shipping at startup pace * Comfort with ambiguity, and the judgement to know when a measurement is good enough to act on ## Description * LLM-as-judge or automated evaluation design, and calibrating it against human judgement * Human annotation programs: rubric authoring, label quality, and annotator throughput as a real constraint * Search ranking, recsys, or online experimentation evaluation - golden-set staleness, offline/online divergence, side-by-side rater agreement. This is the closest existing analog to agentic eval, and it transfers directly * Fine-tuning and evaluating small models: SFT, preference tuning, distillation * Reward modeling, RLHF/RLAIF, or process reward models * Agent trajectory analysis and step-level fault attribution * Prompt engineering as an engineering discipline - versioned, tested, and measured, not tuned by vibes, We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service. ## Related Videos - [Effective Machine Learning - Managing Complexity with MLOps](https://www.wearedevelopers.com/videos/185-effective-machine-learning-managing-complexity-with-mlops) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Same Words, Different Worlds: Who's in Control?](https://www.wearedevelopers.com/videos/2071-same-words-different-worlds-who-s-in-control) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Using LLMs in your Product](https://www.wearedevelopers.com/videos/1186-using-llms-in-your-product) - [How Machine Learning is turning the Automotive Industry upside down](https://www.wearedevelopers.com/videos/61-how-machine-learning-is-turning-the-automotive-industry-upside-down) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Prompt Engineering is a Job of the Past](https://www.wearedevelopers.com/magazine/342-prompt-engineering-is-a-job-of-the-past) - [The Prompt Engineer ✍️](https://www.wearedevelopers.com/magazine/216-the-prompt-engineer) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)