> Markdown version of [/jobs/ext/1862861-ai-ml-evaluation-engineer](https://www.wearedevelopers.com/jobs/ext/1862861-ai-ml-evaluation-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI/ML Evaluation Engineer - **Company:** Booz Allen Hamilton Inc. - **Location:** Atlanta, GA, United States - **Experience:** Expert - **Salary:** $128,700.0 - $292,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, JIRA, Microsoft Azure, Continuous Integration, Information Engineering, Data Governance, Data Systems, Python (Programming Language), Machine Learning, Tensorflow, SQL Databases, Cloud Platform System, Chatbots, Large Language Models, Generative AI, Containerization, Performance Monitor, Machine Learning Operations, Virtual Agents, Data Pipelines - **Published:** July 1, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/87611967/1 ## About the Role * 5+ years of experience with Generative AI, LLMs, AI agents, or RAG applications, and designing, developing, and deploying ML models and AI solutions using Python * 3+ years of experience with AI agents and AI evaluation strategies in enterprise environments * 2+ years of experience with Deep Research evaluation methodologies and AI evaluation workflows * Experience with ML frameworks such as TensorFlow orPyTorch,for production-grade model development * Experience with data engineering usingPySpark, SQL, and Palantir Foundry, including Foundry AIP * Experience withMLOpsplatforms such asMLflowand cloud environments, including Azure * Knowledge of public health, healthcare, or government data systems and associated governance practices * Ability to design andoptimizeAI systems involving retrieval workflows, memory or state management, and real-time decision-support capabilities * Ability to obtain andmaintaina Public Trust or Suitability/Fitness determination based on client requirements * Bachelor's degree in CS, Engineering, or Data Science Nice If You Have: * Experience working in healthcare, biomedical, or government public health AI/ML environments * Experience with conversational AI, chatbot systems, or full-stack AI application development * Experience with containerization, CI/CD, orchestration, and production MLOps pipelines * Experience with Agile delivery environments and tools such as Jira * Experience writing technical documentation and presenting AI/ML solutions to various audiences * Experience integrating enterprise AI tools such as Codex, Claude, or similar AI enablement platforms * Knowledge of enterprise AI governance, compliance, and ethical AI frameworks * Knowledge of AI systems for retrieval, ranking, and scientific evaluation use cases * Ability to collaborate effectively in matrixed, cross-functional organizations * Master's degree in CS, Data Science, ML, or a related field ## Description As an experienced engineer, you know that machine learning (ML) and AI evaluation are critical to understanding and operationalizing massive datasets in support of public health and safety missions. Your ability to evaluate,optimize, and deploy AI-driven systems makes you an integral part of delivering mission-focused solutions for scientists, analysts, and leadership teams. In this role,you'llhelp define and implement scientific AI evaluation and enablement initiatives by translating advanced AI capabilities into practical, mission-specific workflows. You'llcollaborate with a large community of ML engineers, data scientists, architects, and product teams to design scalable ML and Generative AI solutions, including AI agents, retrieval-augmented generation (RAG) pipelines, and enterprise AI evaluation frameworks. You'llapply technicalexpertiseacross AI evaluation, retrieval optimization, memory and state management, and AI agent architecture to support real-time insights and decision-making. Your work will also contribute to enterpriseMLOpscapabilities, data governance standards, and ethical AI practices within regulated public health environments. This positionis located inAtlanta, GA. What You'll Work On: * Build andmaintainscalable data pipelines usingPySparkand Palantir Foundry to support AI, analytics, and scientific evaluation workflows. * Design and implement ML and Generative AI workflows, including AI agents, RAG pipelines, and AI evaluation frameworks. * Integrate advanced AI technologies such as Codex and Claude, to support mission-specific workflows, real-time insights, and decision-making capabilities. * Develop evaluation strategies covering model quality, retrieval optimization, benchmarking, and performance monitoring for deployed AI systems. * Design AI architectures supporting memory, state management, structured and unstructured public health data interaction, and scalable agent orchestration. * Establish data governance, privacy, anonymization, documentation, and ethical AI standards across AI/ML systems and public health data environments. ## Related Videos - [Inside the AI Revolution: How Microsoft is Empowering the World to Achieve More](https://www.wearedevelopers.com/videos/869-inside-the-ai-revolution-how-microsoft-is-empowering-the-world-to-achieve-more) - [Collaboration Quantified: Lessons from Open Source Developer Networks](https://www.wearedevelopers.com/videos/1422-collaboration-quantified-lessons-from-open-source-developer-networks) - [Chatbots are going to destroy infrastructures and your cloud bills](https://www.wearedevelopers.com/videos/1130-chatbots-are-going-to-destroy-infrastructures-and-your-cloud-bills) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Integrate your Cognitive Assistant with 3rd-party DBs and software](https://www.wearedevelopers.com/videos/249-integrate-your-cognitive-assistant-with-3rd-party-dbs-and-software) - [Testing AI Agents: Automated Evaluation for Chatbots & RAG Systems](https://www.wearedevelopers.com/videos/100300-testing-ai-agents-automated-evaluation-for-chatbots-rag-systems) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)