> Markdown version of [/videos/100574-ploomloom-startup-pitch-reliable-evals-for-llm-and-agent-releases-using-plumloom](https://www.wearedevelopers.com/videos/100574-ploomloom-startup-pitch-reliable-evals-for-llm-and-agent-releases-using-plumloom). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Ploomloom - Startup pitch: Reliable evals for LLM and agent releases using Plumloom Stop relying on flaky LLM evaluations for deployment decisions. Plumloom uses independent judges and statistical confidence intervals to automatically block bad AI releases in your pipeline. - **Speakers:** - **Event:** World Congress 2026 North America - **Published:** September 24, 2026 - **Duration:** 3:29 - **URL:** https://www.wearedevelopers.com/videos/100574-ploomloom-startup-pitch-reliable-evals-for-llm-and-agent-releases-using-plumloom ## Access Playback and chapters for this video are available with a Free account. ## Related Moments - [Evaluating LLM applications beyond standard correctness benchmarks](https://www.wearedevelopers.com/videos/100407-red-teaming-your-llm-app-a-hands-on-threat-model-you-can-reuse) (from "Red Teaming Your LLM App -- A Hands-On Threat Model You Can Reuse") - [Introduction to evaluating AI code review agents](https://www.wearedevelopers.com/videos/100421-21-experiments-in-six-weeks-a-playbook-for-improving-your-ai-agent) (from "21 Experiments in Six Weeks: A Playbook for Improving Your AI Agent") - [Breaking complex artificial intelligence evaluations into testable pieces](https://www.wearedevelopers.com/videos/100588-from-boardroom-to-build-pipeline-what-ai-governance-actually-looks-like-in-practice) (from "From Boardroom to Build Pipeline: What AI Governance Actually Looks Like in Practice") - [Programmatic model evaluation and custom metrics via MLflow](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) (from "Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source") - [Executing automated evaluation suites leveraging LLMs as judges](https://www.wearedevelopers.com/videos/1982-stop-guessing-start-measuring-evaluating-rag-systems-with-synthetic-test-data) (from "Stop Guessing, Start Measuring: Evaluating RAG Systems with Synthetic Test Data") - [Evaluating model performance and accuracy using LLM judges](https://www.wearedevelopers.com/videos/100447-from-model-selection-to-smart-routing-how-to-use-the-right-llm-for-every-task) (from "From Model Selection to Smart Routing: How to Use the Right LLM for Every Task") ## Related Articles - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) ## Related Jobs - [Senior AI/ML Engineer](https://www.wearedevelopers.com/jobs/48352-senior-ai-ml-engineer) at **PagerDuty** - [ML Engineer](https://www.wearedevelopers.com/jobs/48448-ml-engineer) at **Docker, Inc.** - [Principal Product Manager, Agentic Evals](https://www.wearedevelopers.com/jobs/48537-principal-product-manager-agentic-evals) at **Expedia** - [LLM Training Engineer](https://www.wearedevelopers.com/jobs/48420-llm-training-engineer) at **Sciforium** - [Staff ML Engineer](https://www.wearedevelopers.com/jobs/48463-staff-ml-engineer) at **Docker, Inc.** - [Partner Sales Director - AI Alliances - Model Providers](https://www.wearedevelopers.com/jobs/48429-partner-sales-director-ai-alliances-model-providers) at **Dynatrace**