World Congress 2026 Europe Jul 10, 2026 Session details

Teaching an LLM to review code … like a Senior Engineer!

Kesha Mykhailov

Are out-of-the-box LLMs sabotaging your pull requests with noisy nitpicks? Discover how one team built an evaluation flywheel to train an AI that reviews code like a senior engineer.

Pause
Mute Enter Fullscreen
#1 about 4 min

Building reliable AI agents for business value

Demonstrable success in automated customer resolution proves that highly effective agents require relentless experimentation and custom models.

#2 about 3 min

Why generic LLMs fail at code reviews

Standard language models produce verbose and irrelevant code feedback because they lack foundational knowledge about unique organizational frameworks.

#3 about 2 min

Leveraging historical context to customize code agents

Automating pull request comments using raw historical design guidelines leads to overwhelming noise without proper deployment testing.

#4 about 2 min

Building an offline evaluation framework for code reviews

Running simulated evaluations inside isolated containers captures structural issues before untested bots can pollute live repositories.

#5 about 3 min

Refining AI prompts through batch evaluations

Using an automated judge to measure severity and correctness helps eliminate false positives triggered by overly rigid prompt formatting.

#6 about 2 min

Gathering human signals to validate agent helpfulness

Mining direct interface feedback and implied commit changes reveals true metrics for an automated reviewer's operational value.

#7 about 3 min

Expanding AI agents to handle pull request approvals

Deconstructing the approval process assigns distinct tasks to specialized sub-agents for catching fundamental bugs and architectural violations.

#8 about 2 min

Designing agent-first interfaces for developer workflows

Constructing specialized interaction protocols ensures that automated assistants can curate evaluation datasets without disrupting natural programming habits.

#9 about 4 min

Preventing deployment risks using deterministic commit limits

Limiting automatic approvals to exceptionally short architectural changes forces developers to embrace continuous delivery and mitigate production flaws.

#10 about 4 min

Shaping sociotechnical systems with character-driven automation

Injecting distinguished personas into pipeline tooling cuts through alert fatigue and naturally encourages engineers to divide dense assignments.

#11 about 2 min

Managing unexpected consequences in automated workflow gating

Intense mechanical bottlenecks frequently cause frustration but inadvertently push engineers to preserve intricate logic rationale inside version histories.

#12 about 3 min

Strategic requirements for implementing enterprise AI agents

Successfully scaling an intelligent code assistant demands continuous background iteration frameworks alongside explicit social reinforcement techniques.

#13 about 4 min

Retaining knowledge sharing inside generative coding environments

Tracking granular design choices during automated coding sessions enables experienced developers to mentor juniors asynchronously.

Matching moments

2:32 min

Reducing pull request cycle times with artificial intelligence

Andrew Boyagi Andrew Boyagi · WWC 2025

1:57 min

Enforcing automated code reviews to manage accelerated delivery

Harald Kirschner · Coffee With Developers

1:31 min

Replacing traditional code reviews with interactive agents

Victor Savkin Victor Savkin · WWC Europe 2026

3:26 min

Introducing LLMs as judges for automated testing

Sebastian Messingfeld Sebastian Messingfeld · WWC Europe 2026

1:31 min

Automating pull request feedback with AI code review agents

Merrill Lutsky Merrill Lutsky · WWC 2025

1:55 min

Shifting developer workloads and realistic AI productivity gains

Chris Heilmann +2 · LIVE

Upcoming sessions on this topic

Open session

World Congress 2026 North America

RoboCoders: Judgment Day: AI-Assisted Engineering Applied - The Battle of Agents

Baruch Sadogursky, Viktor Gamov

Baruch Sadogursky
Viktor Gamov
Open session

World Congress 2026 North America

Agents Can't Iterate Against Tests That Lie

Rocky Warren

Senior Staff Software Engineer at Clipboard

Rocky Warren
Open session

World Congress 2026 North America

DeepAgents: Build Multi-Agent AI Systems That Actually Work

Anagha Rumade, Anjana Umapathy, Apoorva Jaiswal

Anagha Rumade
Anjana Umapathy
Apoorva Jaiswal
Open session

World Congress 2026 North America

Evals Are Infra: Building AI Systems Developers Can Actually Trust

Phoebe Wang

Member of Technical Staff at OpenAI

Phoebe Wang
Open session

World Congress 2026 North America

GitHub’s Team X-Ray: Your Repository Knows More About Your Team Than Your Team Does

Andrea Griffiths

Senior Developer Advocate

Andrea Griffiths
Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong