> Markdown version of [/events/world-congress-2026-north-america/sessions/1959-ploomloom-startup](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1959-ploomloom-startup). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Ploomloom - Startup pitch: Reliable evals for LLM and agent releases using Plumloom - **Date:** Thursday, Sep 24, 2026 - **Time:** 11:50–11:55 (5 min) - **Room:** Outdoor Stage - **Event:** World Congress 2026 North America ## Description AI evals behave like flaky tests. The model under test is probabilistic, and the LLM judging it often is too. The same eval can pass on one run and fail on the next. A single score is one sample, which makes it a shaky basis for a release gate. Plumloom makes the eval score itself more reliable. It calibrates the judging standard, uses multiple independent judges, and for scenario evaluations runs repeated trials to report a confidence interval. That helps separate a real improvement from noise instead of relying on a single number. Autoeval, our open-source CLI with Apache 2.0 licensing and MCP support, brings this into the release path. Bring the scenario, transcript, or OpenTelemetry/OpenInference trace your harness already produces. There is no SDK or observability pipeline to wire up, and CI can gate the release when a score misses your threshold. This pitch covers what an eval really is, why single-run scores can mislead, and how to put reliable, gated evals into your terminal, agent harness, and CI. ## Related talks at this congress - [Evals Are Infra: Building AI Systems Developers Can Actually Trust](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1701-evals-are-infra) — Phoebe Wang - [Your Evals Passed. Your Agent Just Emptied a Database.](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1684-your-evals-passed) — Tejas Pravinbhai Patel - [Taming Rogue Agents: Observability-Driven Evaluation for Production Reliability](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1768-taming-rogue-agents) — Anagha Rumade, Anjana Umapathy, Apoorva Jaiswal - [No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1671-no-single-model-to) — Emmanuel Acheampong ## Watch remotely [Watch live in the app](https://app.wearedevelopers.com)