World Congress 2026 Europe - Virtual Stage Jul 2, 2026 Session details

Stop Guessing, Start Measuring: Evaluating RAG Systems with Synthetic Test Data

Csenge Szabo

Are your RAG applications failing silently? Stop relying on user complaints to catch hallucinations. Learn to independently measure retrieval and generation using RAGAS and synthetic data.

Pause
Mute Enter Fullscreen
#1 about 3 min

Establishing rigorous evaluation pipelines for operational RAG systems

Preventing hallucinated answers and poor data grounding requires setting up comprehensive observability pipelines before application deployment.

#2 about 2 min

Understanding overarching retrieval and generation steps in RAG architectures

Connecting source documents into vector stores supports systemic context retrieval and dynamic model answer generation.

#3 about 2 min

Isolating independent failure surfaces in RAG application pipelines

Diagnosing silent breakages properly involves separating the underlying accuracy of retrieval logic from final component generation.

#4 about 3 min

Overcoming volume and coverage limits in evaluation datasets

Handcrafting exact gold labels manually proves too slow and expensive to accommodate adequate complexity and request variety.

#5 about 4 min

Generating synthetic test data using structured open frameworks

Bypassing generic model traversals in favor of graph-structured tools prevents shallow and overly repetitious question extraction.

#6 about 2 min

Defining varied query types via knowledge graph topologies

Distinguishing between multi-hop and single-hop abstraction ensures evaluation regimes effectively challenge complex chunk retrieval mechanics.

#7 about 3 min

Synthesizing customized test sets from interconnected node relationships

Drawing concrete connections between text snippets through shared themes empowers highly realistic interactive persona emulation.

#8 about 7 min

Automating synthetic test set compilation via Python scripts

Leveraging framework integrations splits raw markdown documentation into optimally overlapping chunks for automated JSON graph formatting.

#9 about 6 min

Diagnosing retrieval and generation pipelines with quantifiable metrics

Measuring context precision alongside noise sensitivity precisely isolates flawed ranking algorithms from poor system prompting.

#10 about 3 min

Executing automated evaluation suites leveraging LLMs as judges

Scoring synthetic dataset queries programmatically against system responses provides an objective baseline for fine-tuning configuration changes.

#11 about 4 min

Incorporating human validation to mitigate automated model bias

Countering inherent system bias requires evaluating synthetically labeled test datasets alongside genuine faults identified in production logs.

#12 about 2 min

Establishing baseline observability and evaluation prior to launch

Adopting quantifiable data metrics from day one proactively prevents system blind spots across evolving architectural revisions.

Matching moments

38 sec

Introducing OpenRAG for custom data pipelines

Phil Nash · Coffee With Developers

1:34 min

Mitigating the inherent challenges of generative AI tools

Mary Grygleski Mary Grygleski · LIVE

3:22 min

Evaluating advanced artificial intelligence platforms for daily recruitment

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

2:12 min

Deploying conversational intelligence tools for complex unstructured data

Niklas Blumenthal Niklas Blumenthal +1 · WWC 2025

1:50 min

Simplifying generative AI deployments using the RagStack opinionated framework

David Leconte David Leconte +1 · WWC 2024

1:44 min

Understanding basic retrieval-augmented generation architectures in chatbots

Stan Girard Stan Girard · WWC 2024

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Who Tests the AI? Building Trustworthy AI Systems at Enterprise Scale

Him Raj Singh

PayPal, Manager, Software Engineer

Him Raj Singh
Open session

World Congress 2026 North America

Evals Are Infra: Building AI Systems Developers Can Actually Trust

Phoebe Wang

Member of Technical Staff at OpenAI

Phoebe Wang
Open session

World Congress 2026 North America

Your Evals Passed. Your Agent Just Emptied a Database.

Tejas Pravinbhai Patel

IEEE Award-Winning Researcher | Best Keynote Speaker | Sr. Software Engineer at Amazon | AI Systems & Agent Architect

Tejas Pravinbhai Patel
Open session

World Congress 2026 North America

Reinventing Testing Practices in the AI Era

Eric Deandrea

Java Champion & Senior Principal Software Engineer, IBM

Eric Deandrea
Open session

World Congress 2026 North America

Closing the Visibility Gap: Lessons from Safety Critical Agentic Systems

Vivek Pandit

Principal Engineer at Cadence

Vivek Pandit
Open session

World Congress 2026 North America

RoboCoders: Judgment Day: AI-Assisted Engineering Applied - The Battle of Agents

Baruch Sadogursky, Viktor Gamov

Baruch Sadogursky
Viktor Gamov