World Congress 2025

Beyond the Prompt: Evaluating, Testing, and Securing LLM Applications

July 10, 2025 14:10 – 14:40 · 30 min Stage 2
ai llms automation cybersecurity testing

What this session covers

When you change prompts or modify the Retrieval-Augmented Generation (RAG) pipeline in your LLM applications, how do you know it’s making a difference? You don’t—until you measure. But what should you measure, and how? Similarly, how can you ensure your LLM app is resilient against prompt injections or avoids providing harmful responses? More robust guardrails on inputs and outputs are needed beyond basic safety settings.

In this talk, we’ll explore various evaluation frameworks such as Vertex AI Evaluation, DeepEval, and Promptfoo to assess LLM outputs, understand the types of metrics they offer, and how these metrics are useful. We’ll also dive into testing and security frameworks like LLM Guard to ensure your LLM apps are safe and limited to precisely what you need.

Related talks at this congress

Open session

World Congress 2025

July 11, 2025 · 13:00–13:30

Stage 4

Prompt Injection, Poisoning & More: The Dark Side of LLMs

Keno Dreßel

Principal Consultant & Head of AI @ SQUER

Keno Dreßel
Open session

World Congress 2025

July 10, 2025 · 12:10–12:40

Stage 6 - Red Hat

Beyond the Hype: Building Trustworthy and Reliable LLM Applications with Guardrails

Alex Soto

Director of Developer Experience at Red Hat

Alex Soto
Open session

World Congress 2025

July 10, 2025 · 15:30–16:00

Stage 2

The Limits of Prompting: ArchitectingTrustworthy Coding Agents

Nimrod Kor

Co-founder & CTO @ Baz

Nimrod Kor
Open session

World Congress 2025

July 9, 2025 · 11:00–13:00

M4 (40 Seats)

A Jumpstart to Using LLMs for Test Design

Pierre Baum, Rahul

Pierre Baum
Rahul
All sessions at this congress