> Markdown version of [/jobs/ext/1345549-ai-evaluation-engineer](https://www.wearedevelopers.com/jobs/ext/1345549-ai-evaluation-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Evaluation Engineer - **Company:** GovWorx, Inc. - **Location:** United States (Remote available) - **Experience:** Experienced - **Salary:** $120,000.0 - $170,000.0 - **Contract:** Permanent contract - **Skills:** A/B Testing, Artificial Intelligence, Amazon S3, Data Analysis, Database Queries, Statistical Hypothesis Testing, Python (Programming Language), Machine Learning, Performance Tuning, Power BI, Software Engineering, SQL Databases, Tableau (Software), Large Language Models, Prompt Engineering, Generative AI, Machine Learning Operations - **Published:** July 19, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=fa921d016588b93c ## About the Role Clearance: Must have US citizenship and pass FBI fingerprint and background check in multiple states, * Must have US citizenship and pass FBI fingerprint and background check in multiple states * 3+ years of experience in software engineering, machine learning, data science, or a related technical field * Experience designing evaluation metrics and interpreting AI model performance * Understanding of statistical methods including hypothesis testing and experiment design * Strong Python development experience * Strong SQL skills with experience analyzing large datasets * Experience building or supporting production LLM or Generative AI applications * Experience with prompt engineering and systematic prompt evaluation, * Experience using AI evaluation or observability platforms such as Langfuse, LangSmith, MLflow, or Label Studio * Experience with AWS services such as Bedrock, Lambda, S3, Glue, or SageMaker * Experience building dashboards using Tableau, Sisense, Power BI, or similar tools * Knowledge of Responsible AI principles and evaluation methodologies ## Description We're looking for an experienced AI Evaluation Engineer to help build and improve the next generation of AI systems used by public safety agencies across the country. This role sits at the intersection of AI engineering, prompt engineering, and data science. You'll own the evaluation and continuous improvement of production AI systems, developing automated evaluation pipelines, designing prompt experiments, analyzing model performance, and building tooling that enables rapid iteration. You'll work closely with data scientists, data engineers, and product managers to ensure our AI systems remain accurate, reliable, and trustworthy in real-world public safety environments., * Design, build, and maintain automated AI evaluation pipelines for production LLM applications * Develop prompt engineering strategies and iterate on prompts and compare LLMs using quantitative evaluation methods * Build offline evaluation datasets and regression testing frameworks to measure AI performance over time * Analyze production AI behavior using Python, SQL, and statistical techniques to identify opportunities for improvement * Design experiments, A/B tests, and benchmarking methodologies for evaluating prompt and model changes * Develop dashboards and reporting that communicate AI quality, reliability, and performance metrics * Partner with engineering and product teams to safely deploy and monitor improvements to production AI systems * Investigate model failures through detailed error analysis and recommend improvements to prompts, evaluation datasets, and workflows * Help establish best practices for Responsible AI, evaluation methodologies, and continuous model improvement ## Related Videos - [Bringing the power of AI to your application.](https://www.wearedevelopers.com/videos/1010-bringing-the-power-of-ai-to-your-application) - [WeAreDevelopers LIVE - CSS is DOOMed](https://www.wearedevelopers.com/videos/1838-wearedevelopers-live-css-is-doomed) - [Beyond Dashboards: Fixing Text-to-SQL with Semantic RAG](https://www.wearedevelopers.com/videos/2036-beyond-dashboards-fixing-text-to-sql-with-semantic-rag) - [Evals vs. Evil - AI and Package Security - Laurie Voss](https://www.wearedevelopers.com/videos/2131-evals-vs-evil-ai-and-package-security-laurie-voss) - [REST, GraphQL, gRPC, and more: A comparison of modern API styles](https://www.wearedevelopers.com/videos/100247-rest-graphql-grpc-and-more-a-comparison-of-modern-api-styles) - [Beyond the Benchmark: How to Evaluate AI Agents in the Real World](https://www.wearedevelopers.com/videos/100269-beyond-the-benchmark-how-to-evaluate-ai-agents-in-the-real-world) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [The State of WebDev AI 2025 Results: What Can We Learn?](https://www.wearedevelopers.com/magazine/581-the-state-of-webdev-ai-2025-results-what-can-we-learn) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)