World Congress 2026 Europe • Jul 10, 2026 • Session details

Teaching an LLM to review code … like a Senior Engineer!

Kesha Mykhailov

Are out-of-the-box LLMs sabotaging your pull requests with noisy nitpicks? Discover how one team built an evaluation flywheel to train an AI that reviews code like a senior engineer.

Pause
Mute Enter Fullscreen
#1 about 4 min

Building reliable AI agents for business value

Demonstrable success in automated customer resolution proves that highly effective agents require relentless experimentation and custom models.

#2 about 3 min

Why generic LLMs fail at code reviews

Standard language models produce verbose and irrelevant code feedback because they lack foundational knowledge about unique organizational frameworks.

#3 about 2 min

Leveraging historical context to customize code agents

Automating pull request comments using raw historical design guidelines leads to overwhelming noise without proper deployment testing.

#4 about 2 min

Building an offline evaluation framework for code reviews

Running simulated evaluations inside isolated containers captures structural issues before untested bots can pollute live repositories.

#5 about 3 min

Refining AI prompts through batch evaluations

Using an automated judge to measure severity and correctness helps eliminate false positives triggered by overly rigid prompt formatting.

#6 about 2 min

Gathering human signals to validate agent helpfulness

Mining direct interface feedback and implied commit changes reveals true metrics for an automated reviewer's operational value.

#7 about 3 min

Expanding AI agents to handle pull request approvals

Deconstructing the approval process assigns distinct tasks to specialized sub-agents for catching fundamental bugs and architectural violations.

#8 about 2 min

Designing agent-first interfaces for developer workflows

Constructing specialized interaction protocols ensures that automated assistants can curate evaluation datasets without disrupting natural programming habits.

#9 about 4 min

Preventing deployment risks using deterministic commit limits

Limiting automatic approvals to exceptionally short architectural changes forces developers to embrace continuous delivery and mitigate production flaws.

#10 about 4 min

Shaping sociotechnical systems with character-driven automation

Injecting distinguished personas into pipeline tooling cuts through alert fatigue and naturally encourages engineers to divide dense assignments.

#11 about 2 min

Managing unexpected consequences in automated workflow gating

Intense mechanical bottlenecks frequently cause frustration but inadvertently push engineers to preserve intricate logic rationale inside version histories.

#12 about 3 min

Strategic requirements for implementing enterprise AI agents

Successfully scaling an intelligent code assistant demands continuous background iteration frameworks alongside explicit social reinforcement techniques.

#13 about 4 min

Retaining knowledge sharing inside generative coding environments

Tracking granular design choices during automated coding sessions enables experienced developers to mentor juniors asynchronously.

Matching moments

2:32 min

Reducing pull request cycle times with artificial intelligence

Andrew Boyagi Andrew Boyagi · World Congress 2025

1:57 min

Enforcing automated code reviews to manage accelerated delivery

Harald Kirschner Harald Kirschner · Coffee With Developers

1:31 min

Replacing traditional code reviews with interactive agents

Victor Savkin Victor Savkin · World Congress 2026 Europe

3:26 min

Introducing LLMs as judges for automated testing

Sebastian Messingfeld Sebastian Messingfeld · World Congress 2026 Europe

1:31 min

Automating pull request feedback with AI code review agents

Merrill Lutsky Merrill Lutsky · World Congress 2025

1:55 min

Shifting developer workloads and realistic AI productivity gains

Chris Heilmann Chris Heilmann +2 · LIVE