> Markdown version of [/jobs/ext/2156751-sr-evaluation-engineer](https://www.wearedevelopers.com/jobs/ext/2156751-sr-evaluation-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Evaluation Engineer - **Company:** LogicMonitor, Inc. - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $158,400.0 - $217,800.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Continuous Integration, Python (Programming Language), Machine Learning, Regression Analysis, Regression Testing, Software Engineering, Retrieval-Augmented Generation, Large Language Models, Multi-Agent Systems, Prompt Engineering, Model Validation, Machine Learning Operations - **Published:** August 20, 2026 - **Apply:** https://dejobs.org/x/x/2B11F8F3C44F470B91685FE923F47B34/job/ ## About the Role * 5+ years of experience in software engineering, machine learning, applied AI, or a related field. * Strong Python engineering skills and experience building production systems. * Hands-on experience with AI evaluation, experimentation, testing, and quality frameworks. * Experience using multiple LLM and agent evaluation frameworks, such as LangSmith, Arize Phoenix, Braintrust, DeepEval, Ragas, TruLens, OpenAI Evals, MLflow, or comparable platforms. * Ability to select, customize, and integrate evaluation frameworks for offline testing, online monitoring, regression analysis, experimentation, model and prompt comparison, and release gating. * Strong understanding of LLMs, agents, retrieval-augmented generation, prompt engineering, tool calling, and context engineering. * Experience evaluating non-deterministic, multi-step, or multi-agent AI systems. * Ability to translate human and domain-expert judgment into test cases, evaluation rubrics, scoring functions, and automated graders. * Experience with LLM-as-a-judge techniques, including grader design, calibration, reliability measurement, and alignment with expert human judgment. * Experience with regression testing, CI/CD, production monitoring, behavioral drift detection, and failure analysis. * Strong analytical, systems-thinking, and communication skills., At this time, we are able to consider candidates who are authorized to work in the United States on a full-time, permanent basis without requiring new or initial employer-sponsored work authorization. Candidates who currently hold valid U.S. work authorization that can be transferred to a new employer (such as certain H-1B statuses) may be considered on a case-by-case basis. ## Description * Define quality metrics for incident diagnostics, root-cause analysis, alert correlation, grounding, tool use, safety, and operational usefulness. * Build offline and online evaluation pipelines in Python and integrate them with CI/CD, experimentation, model selection, prompt iteration, and release gating. * Lead the creation and maintenance of golden datasets and regression suites using alerts, events, metrics, logs, traces, topology, configuration data, incident timelines, change records, ITSM workflows, and historical investigation outcomes. * Build representative, customer-specific scenarios covering different technologies, failure modes, operational patterns, and environmental constraints. * Use human-authored and AI-assisted methods to generate regression, edge, adversarial, rare, ambiguous, and incomplete-context test cases. * Treat evaluation datasets and test suites as first-class components that evolve alongside Edwin AI. * Design step-level and trajectory-level evaluations for multi-step and multi-agent workflows. * Assess both final outcomes and intermediate behavior, including planning, reasoning consistency, retrieval, evidence use, tool selection, tool parameters, state transitions, escalation decisions, and human-in-the-loop approvals. * Identify whether failures originate from models, prompts, retrieval, data quality, tools, agent logic, orchestration, or infrastructure. * Evaluate capabilities including incident investigation, on-call assistance, operational question answering, change-impact analysis, remediation recommendations, automated resolution, infrastructure operations, and ITSM and observability integrations. * Design and calibrate LLM-based graders against expert human judgment. * Monitor AI quality and behavioral drift in production, and convert failures and customer feedback into new tests and safeguards. * Establish evaluation-driven development practices and mentor other engineers. ## Related Videos - [Why LLMs Need Observability and How to Do It](https://www.wearedevelopers.com/videos/2117-why-llms-need-observability-and-how-to-do-it) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Introduction to Responsible AI: Balancing Value and Risk](https://www.wearedevelopers.com/videos/1972-introduction-to-responsible-ai-balancing-value-and-risk) - [Trunk-Based Development at Scale: Real-World Insights from a High-Traffic Luxury E-Commerce Platform](https://www.wearedevelopers.com/videos/1435-trunk-based-development-at-scale-real-world-insights-from-a-high-traffic-luxury-e-commerce-platform) - [You are not my model anymore - understanding LLM model behavior](https://www.wearedevelopers.com/videos/1464-you-are-not-my-model-anymore-understanding-llm-model-behavior) - [Evals vs. Evil - AI and Package Security - Laurie Voss](https://www.wearedevelopers.com/videos/2131-evals-vs-evil-ai-and-package-security-laurie-voss) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Never delegate the understanding](https://www.wearedevelopers.com/magazine/749-never-delegate-the-understanding)