World Congress 2025

The Limits of Prompting: ArchitectingTrustworthy Coding Agents

July 10, 2025 15:30 – 16:00 · 30 min Stage 2

What this session covers

LLMs can summarize, autocomplete, and comment on code, but when applied to tasks like code review, they often fall apart. We’ve run agents across real PRs in production environments and observed consistent failure modes: overconfidence, repetition, and lack of judgment. This talk examines the architectural reasons behind that breakdown—why transformer-based models struggle with intent, structure, and impact when reviewing code - and how reasoning phases, memory, and reflection improve both precision and trust. Using real-world examples from Baz’s code review agent, we’ll show why prompt engineering isn’t enough, and what it takes to move from generic pattern-matching to useful, reliable feedback in large, evolving codebases. Expect concrete takeaways for anyone building or evaluating LLM-based developer tools.

Related talks at this congress

Open session

World Congress 2025

July 11, 2025 · 13:40–13:50

Airstream 1

Prompts are boring. Also, avoidable.

Marcin Koralewski

Marcin Koralewski, Lead Data Scientist at Zoovu

Marcin Koralewski
Open session

World Congress 2025

July 11, 2025 · 16:20–16:50

Stage 2

How we built an AI-powered code reviewer in 80 hours

Yan Cui

Developer Advocate at Lumigo

Yan Cui
Open session

World Congress 2025

July 10, 2025 · 12:10–12:40

Stage 6 - Red Hat

Beyond the Hype: Building Trustworthy and Reliable LLM Applications with Guardrails

Alex Soto

Director of Developer Experience at Red Hat

Alex Soto
Open session

World Congress 2025

July 10, 2025 · 14:10–14:40

Stage 2

Beyond the Prompt: Evaluating, Testing, and Securing LLM Applications

Mete Atamel

Software Engineer and Developer Advocate at Google

Mete Atamel
All sessions at this congress