> Markdown version of [/jobs/ext/1182056-ai-engineer-evaluation](https://www.wearedevelopers.com/jobs/ext/1182056-ai-engineer-evaluation). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Engineer, Evaluation - **Company:** Distyl AI - **Location:** New York, NY, United States - **Experience:** Experienced - **Salary:** $150,000.0 - $250,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Software Debugging, Python (Programming Language), Systems Development Life Cycle, Software Engineering, Large Language Models, Prompt Engineering - **Published:** July 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=a13b9bfe43c1c05d ## About the Role * 2+ years of software engineering experience * Strong Python Engineering Skills: Write clean, maintainable Python and are comfortable building evaluation and experimentation pipelines that run in production environments. You treat evaluation code with the same rigor as application code * Experience with Evaluation-Driven or Experiment-Driven Development: Experience using structured evaluation or experimentation frameworks to drive system iteration, and understand the pitfalls of overfitting to metrics that don't reflect real outcomes * Ability to Translate Human Judgment into Code: Work with subject matter experts to elicit high-quality judgments and encode them into test cases, scoring functions, and graders that scale * Systems-Oriented Mindset: Understand how evaluation interacts with prompts, agents, data, and deployment. You design evaluation systems that support fast iteration while maintaining trust and safety in production * AI-Native Working Style: Use AI tools to generate tests, analyze failures, explore edge cases, and accelerate debugging and iteration * Travel: Travel between 10-50% of the time, depending on the project, your role and level of interest in doing so ## Description * Design and implement evaluation frameworks that enable Evaluation-Driven Development for AI systems deployed in customer environments * Define how system quality is measured in each domain, ensuring that evaluation signals reflect real user needs, domain constraints, and business objectives * Build and maintain golden test cases and regression suites in Python, using both human-authored and AI-assisted test generation to capture critical behaviors and edge cases. These test suites are treated as first-class system components that evolve alongside the AI system itself * Develop and maintain evaluation pipelines-offline and online-that integrate directly into system iteration loops. Evaluation results inform prompt design, agent logic, model selection, and release readiness, ensuring that system changes are driven by measurable improvements rather than intuition alone * Define, calibrate, and operate LLM-based graders, aligning automated judgments with expert human assessments. They investigate where evaluation signals diverge from real-world outcomes and refine grading approaches to maintain signal quality as systems and domains evolve * Work closely with Forward Deployed AI Engineers, Architects, Product Engineers, AI Strategists, and domain experts to ensure evaluation frameworks meaningfully guide system development and deployment in production ## Related Videos - [Are We All Prompt Engineers? How AI Changed What It Means to Build Software](https://www.wearedevelopers.com/videos/1980-are-we-all-prompt-engineers-how-ai-changed-what-it-means-to-build-software) - [WeAreDevelopers LIVE - Markdown, Liquid and Checkouts](https://www.wearedevelopers.com/videos/1814-wearedevelopers-live-markdown-liquid-and-checkouts) - [What I learned as a developer from accidents in space](https://www.wearedevelopers.com/videos/642-what-i-learned-as-a-developer-from-accidents-in-space) - [Official Opening of WeAreDevelopers World Congress 2026](https://www.wearedevelopers.com/videos/100000-official-opening-of-wearedevelopers-world-congress-2026) - [The Software Engineer 2030: From Coder To AI Orchestrator? - Patrick Schnell](https://www.wearedevelopers.com/videos/1825-the-software-engineer-2030-from-coder-to-ai-orchestrator-patrick-schnell) - [Engineering Productivity: Cutting Through the AI Noise](https://www.wearedevelopers.com/videos/1691-engineering-productivity-cutting-through-the-ai-noise) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [The State of WebDev AI 2025 Results: What Can We Learn?](https://www.wearedevelopers.com/magazine/581-the-state-of-webdev-ai-2025-results-what-can-we-learn) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Transforming Software Development: The Role of AI and Developer Tools](https://www.wearedevelopers.com/magazine/527-transforming-software-development-the-role-of-ai-and-developer-tools)