> Markdown version of [/jobs/ext/2709009-backend-engineer-ai-evaluations](https://www.wearedevelopers.com/jobs/ext/2709009-backend-engineer-ai-evaluations). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Backend Engineer, AI Evaluations - **Company:** Ellipsis Health, Inc. - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $160,000.0 - $210,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), JavaScript (Programming Language), Application Programming Interfaces (APIs), Artificial Intelligence, Automation of Tests, Configuration Management, Continuous Integration, Data Integration, Software Debugging, Programming Tools, EHealth, Python (Programming Language), Software Engineering, TypeScript, Data Logging, Chatbots, Large Language Models, Grafana, Backend, Tools for Reporting, Data Pipelines - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-backend-engineer-ai-evaluations-ellipsis-health-8344048 ## About the Role * 5+ years of professional software engineering experience, with a strong focus on building backend systems, platforms, or developer tooling. * Proven experience designing and maintaining production-grade infrastructure with code, including APIs, services, and data pipelines. * Strong proficiency in at least one general-purpose programming language (e.g., Python, Typescript/Javascript, Java, or similar). * Experience using test automation frameworks, evaluation pipelines, or CI/CD-integrated testing systems. * Familiarity with observability and debugging tools (logging, metrics, tracing) and building internal tools that improve developer and QA workflows. * Strong debugging skills and a methodical approach to diagnosing production and evaluation issues. * Ability to collaborate effectively across engineering, QA, and operations teams, translating requirements into reliable, maintainable systems. * Product-minded approach to infrastructure, with attention to usability, documentation, and long-term maintainability. Preferred: * Experience working with complex, multi-component systems (e.g., ASR, LLMs, TTS, or other ML-powered services) * Experience working in healthcare or other regulated environments, including awareness of HIPAA and PHI handling. * Familiarity with conversational AI or voice agents, including multi-turn dialogue, latency constraints, and error recovery. * Familiarity with LLM observability or evaluation tools (e.g., Langfuse, prompt eval frameworks). * Background in digital health, care coordination, or patient-facing systems. ## Description We're seeking a strong Senior Backend Engineer to join our AI Evaluation team, focused on building the infrastructure and internal tooling that enable reliable, repeatable evaluation of AI systems in production. In this role, you'll develop the scaffolding around our evaluations platform - API/data integrations, test configuration management, and reporting tools that make evaluations easy to run, extend, and operate at scale. You'll also build and maintain evaluation frameworks for individual system components such as ASR, LLMs, TTS, knowledge bases, and guardrails, productionalizing these frameworks to support regular analysis, regression detection, and continuous monitoring. In addition, you'll create debugging tools that make it easy to inspect end-to-end calls, trace failures, and surface all relevant signals in one place - empowering internal teams to diagnose issues quickly and confidently. The ideal candidate is a pragmatic, infrastructure-minded engineer who enjoys turning ad hoc analysis into durable systems, cares deeply about developer experience, holds a high bar for software-quality and takes pride in building tooling that makes complex AI systems observable, testable, and easier to operate., * Build and maintain infrastructure and tooling for the AI evaluations platform used by internal teams, including automated testing platform for AI voice agents, debugging and observability tools. * Develop and productionalize evaluation frameworks for individual system components such as ASR, LLMs, TTS, knowledge bases, and guardrails. * Partner with ML, engineering and QA teams to translate evaluation requirements into robust, maintainable infrastructure and tooling. * Improve developer experience by making evaluation systems easy to extend, well-documented, and reliable in day-to-day use. * Ensure evaluation tooling meets production standards for reliability, performance, and maintainability. ## Related Videos - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [How E.On productionizes its AI model & Implementation of Secure Generative AI.](https://www.wearedevelopers.com/videos/623-how-e-on-productionizes-its-ai-model-implementation-of-secure-generative-ai) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Nest.js - TypeScript in the backend can also be clean](https://www.wearedevelopers.com/videos/1033-nest-js-typescript-in-the-backend-can-also-be-clean) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)