World Congress 2026 North America • Sep 24, 2026 • Session details

Ploomloom - Startup pitch: Reliable evals for LLM and agent releases using Plumloom

Stop relying on flaky LLM evaluations for deployment decisions. Plumloom uses independent judges and statistical confidence intervals to automatically block bad AI releases in your pipeline.

Ploomloom - Startup pitch: Reliable evals for LLM and agent releases using Plumloom thumbnail

Checking access…

Playback and chapters load privately for Free videos.

Matching moments

2:36 min

Evaluating LLM applications beyond standard correctness benchmarks

Saloni Garg Saloni Garg · World Congress 2026 North America

1:56 min

Introduction to evaluating AI code review agents

Sofia Rest Sofia Rest · World Congress 2026 North America

1:31 min

Breaking complex artificial intelligence evaluations into testable pieces

Anita Ganti Anita Ganti · World Congress 2026 North America

4:42 min

Programmatic model evaluation and custom metrics via MLflow

Viktoria Semaan Viktoria Semaan · World Congress 2026 Europe

2:49 min

Executing automated evaluation suites leveraging LLMs as judges

Csenge Szabo Csenge Szabo · Europe 2026 Virtual

4:04 min

Evaluating model performance and accuracy using LLM judges

Viktoria Semaan Viktoria Semaan · World Congress 2026 North America