World Congress 2025 Aug 20, 2025 Session details

The Limits of Prompting: ArchitectingTrustworthy Coding Agents

Nimrod Kor

Prompting is insufficient for complex code evaluation. Stop relying on fragile, monolithic AI agents. Learn to architect trustworthy, multi-agent systems that genuinely understand your codebase.

Pause
Mute Enter Fullscreen
#1 about 1 min

Automating pull requests with a code review agent

The review agent processes codebase details to generate module summaries and enforce programming rules.

#2 about 2 min

Building a prototype repository listener and wrapper

Creating a listener wrapper aggregates commit metrics for seamless structured prompt communication.

#3 about 2 min

Analyzing semantic context in early prototype testing

Deep language logic accurately detects misleading variable naming conventions regardless of underlying syntax configurations.

#4 about 2 min

Iterating on prompt engineering for edge cases

Advanced prompt modifications gracefully circumvent edge cases like outdated model knowledge and misleading developer comments.

#5 about 2 min

Evaluating the pros and cons of benchmarking

Consistent output benchmarking directly contrasts improved code quality tests against actual token execution expenses.

#6 about 2 min

Best practices for maintaining dynamic evaluation datasets

Scaling functional code evaluation datasets continuously prevents static decay during structured output model validation.

#7 about 3 min

Implementing a concurrent benchmarking and tracking pipeline

A comprehensive automated evaluation pipeline accurately maps execution hit rates and tracks output noise ratios.

#8 about 3 min

Splitting multi-step tasks across specialized agents

Separating complex analysis commands into modular independent agents significantly prevents contextual logic hallucinations.

#9 about 2 min

Integrating repository concepts via retrieval augmented generation

Retrieval augmented generation surfaces highly relevant architectural similarities based on distinct historical code embeddings.

#10 about 2 min

Traversing module connections with abstract syntax trees

Parsing abstract syntax tree relationships directly maps codebase structures to automatically prevent component duplication.

#11 about 2 min

Expanding model scope with ticket and runtime data

Linking external ticket parameters and deployment scopes improves code feedback realism regarding system application performance.

#12 about 1 min

Benchmarking the fully contextualized code review agent

Gathering comprehensive environmental data immediately elevates model verification rates without invoking inefficient prompt loops.

#13 about 2 min

Generating custom guidelines from open source repositories

Compiling feedback from major open source entities establishes practical development rules for common integration environments.

#14 about 2 min

Key takeaways for architecting reliable software agents

Rethinking autonomous software logic through stringent testing and deep context layers achieves reliable engineering automation.

Matching moments

1:31 min

Replacing traditional code reviews with interactive agents

Victor Savkin Victor Savkin · World Congress 2026 Europe

2:06 min

Rethinking team structures around AI agent capabilities

Mike Mike · World Congress 2025

1:55 min

Shifting developer workloads and realistic AI productivity gains

Chris Heilmann +2 · LIVE

3:23 min

Integrating intent-based code generation and agent implementation

1:38 min

Evaluating agent code via previews and critic models

Guillaume Moigneu Guillaume Moigneu · World Congress 2026 Europe

2:34 min

Balancing developer autonomy with the adoption of coding agents

Clemens Wasner Clemens Wasner +4 · World Congress 2026 Europe

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 24, 2026 · 14:25–14:35

Outdoor Stage

SecurePrompt: Building a Pre-Flight Security Layer for Agentic AI

Ravi Sastry Kadali

AI/ML Engineer at General Motors

Ravi Sastry Kadali
Open session

World Congress 2026 North America

September 24, 2026 · 10:20–10:50

Stage 6

Practices, Not Prompts: Scale GitHub Copilot with AI-Ready Repositories

Luis Pujols

Staff Customer Success Architect, GitHub

Luis Pujols
Open session

World Congress 2026 North America

September 25, 2026 · 13:30–14:00

Stage 6

The Autonomous Pull Request: Let Agents Ship Without Surrendering Control

Sam Jarvinen

Senior Solutions Engineer, GitHub

Sam Jarvinen
Open session

World Congress 2026 North America

September 24, 2026 · 13:30–15:30

Stage 8

The Developer's Guide to an Agentic Git Forge

Lizzie Siegle, Evis Drenova

Lizzie Siegle
Evis Drenova
Open session

World Congress 2026 North America

September 24, 2026 · 15:30–16:00

Stage 5

The reviewer can't be the author: independent verification for AI-generated code

Manish Kapur

VP, Product and Solutions at Sonar

Manish Kapur
Open session

World Congress 2026 North America

September 24, 2026 · 16:10–16:40

Stage 1

Sandboxing the Swarm: Building Secure, Serverless AI Agents with Wasm

Lena Hall, Thorsten Hans

Lena Hall
Thorsten Hans