World Congress 2026 Europe Jul 10, 2026 Session details

LLMs in the wild: Building an AI agent that survives production

Steven Mi , Giampaolo Casolla

Stop letting prompt tweaks break your production AI. Learn how GetYourGuide used test-driven development and rigorous CI/CD evaluations to safely deploy over 150 prompt modifications.

Pause
Mute Enter Fullscreen
#1 about 3 min

Navigating the gap between prototype and production

Exploring the common challenge of migrating language model prompts to predictable behaviors in production systems.

#2 about 3 min

Motivation for introducing a conversational discovery agent

How changes in the search landscape and user preferences drove the need for automated travel discovery.

#3 about 2 min

Extracting complex semantic intent from user queries

Transforming unstructured and highly constrained travel plans into parseable criteria for backend services.

#4 about 3 min

Moving from API wrappers to agentic orchestration

Orchestrating states locally to improve performance and prevent data leakage observed in platform API wrappers.

#5 about 6 min

Architecting the multi-step travel discovery pipeline

Breaking the discovery journey into specific intent, planning, retrieval, relevance, and drafting components.

#6 about 2 min

Optimizing inference latency and node failure resilience

Implementing fan-out parallelization, fallback paths, and varying model complexities to reduce processing overhead.

#7 about 4 min

Preventing feature regression using test-driven development

How continuous prompt iterations lead to edge case failures without a rigorous evaluation framework.

#8 about 3 min

Designing concrete metrics for deterministic component evaluations

Crafting handcrafted input targets specifying expected fields, explicit null values, and routing logic flags.

#9 about 4 min

Automating regression checks in continuous integration loops

Triggering automated tests with a pass-at-k strategy to confidently merge iterations of variable models.

#10 about 2 min

Key takeaways for reliable production agent integrations

Integrating rigorous deterministic tests from inception and prioritizing system-grade development for feature deployments.

#11 about 2 min

Tooling and language coverage in question responses

Discussing the application of specific observation frameworks and processing inputs in numerous global languages.

Matching moments

3:26 min

Introducing LLMs as judges for automated testing

Sebastian Messingfeld Sebastian Messingfeld · World Congress 2026 Europe

2:21 min

Transitioning from AI co-pilots to AI-native products

Jordan Tigani Jordan Tigani · World Congress 2026 Europe

1:12 min

Enhancing conversational intent through modern large language models

Nathaniel Okenwa Nathaniel Okenwa · World Congress 2025

1:36 min

Differentiating AI search layers from standard LLMs

Klaus-M. Schremser Klaus-M. Schremser · World Congress 2026 Europe

5:30 min

Building components of a real-world LLM lifecycle

Maxim Salnikov Maxim Salnikov · LIVE

1:39 min

Leveraging generative AI and agents for executive productivity

Katrin Lehmann Katrin Lehmann +1 · Coffee With Developers

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 24, 2026 · 17:30–18:00

Stage 6

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

September 23, 2026 · 13:15–15:15

Stage 9

DeepAgents: Build Multi-Agent AI Systems That Actually Work

Anagha Rumade, Anjana Umapathy, Apoorva Jaiswal

Anagha Rumade
Anjana Umapathy
Apoorva Jaiswal
Open session

World Congress 2026 North America

September 24, 2026 · 11:20–11:25

Outdoor Stage

Finding the Edges: Testing, Evaluating, and Monitoring Voice AI Agents Before Your Users Do

Matt Wyman

CEO of Okareo

Matt Wyman
Open session

World Congress 2026 North America

September 24, 2026 · 11:40–12:10

Stage 4

Taming Rogue Agents: Observability-Driven Evaluation for Production Reliability

Anagha Rumade, Anjana Umapathy, Apoorva Jaiswal

Anagha Rumade
Anjana Umapathy
Apoorva Jaiswal
Open session

World Congress 2026 North America

September 23, 2026 · 10:45–12:45

Stage 10

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

September 25, 2026 · 15:45–15:55

Outdoor Stage

Closing the Visibility Gap: Lessons from Safety Critical Agentic Systems

Vivek Pandit

Frontier AI Lead at Turing

Vivek Pandit