World Congress 2026 Europe

11 Principles for Evaluating AI Dev Tools

July 10, 2026 14:20 – 14:50 · 30 min Stage 11

What this session covers

Benchmarks measure narrow capabilities. Demos show best-case scenarios. Neither tells you whether AI-generated code will survive production or whether that shiny new tool deserves a place in your stack.

This talk presents a unified framework of 11 principles for evaluating AI-generated code and the tools that manage it. Because code quality and tool quality are inseparable: bad tools generate bad code, and bad code evaluation processes never catch it.

We’ll reframe AI use with responsibility boundaries where the core questions shift from “is it fast?” to “can this be understood under pressure, safely changed, and defended to a stakeholder?”

Through real-world patterns from teams adopting AI across their SDLC, we’ll apply these principles to distinguish tools that surface risk from tools that hide it.

Attendees will leave with a practical rubric to decide which AI tools to trust, which to constrain, and how to keep human judgment at the center of fast-moving, AI-augmented engineering.

Related talks at this congress

Open session

World Congress 2026 Europe

July 9, 2026 · 13:00–15:00

Room M6 (40 Seats)

High Quality AI Development

David Tielke

Consultant, Coach and Trainer at www.David-Tielke.de

David Tielke
Open session

World Congress 2026 Europe

July 10, 2026 · 13:00–13:30

Stage 8 - powered by Red Hat

Beyond the Benchmark: How to Evaluate AI Agents in the Real World

Taylor Jordan Smith

Senior AI Developer Advocate at Red Hat

Taylor Jordan Smith
Open session

World Congress 2026 Europe

July 9, 2026 · 13:00–15:00

Room M5 (18 Seats)

Maintainable and Testable Code in the Age of AI

Dennis Doomen

Principal Consultant at Aviva Solutions

Dennis Doomen
Open session

World Congress 2026 Europe

July 9, 2026 · 10:50–11:20

Stage 9

Back to the Roots: Testing in the Age of AI

Jakub Janczyk

Senior Engineer at Kit

Jakub Janczyk
All sessions at this congress