World Congress 2026 Europe Jul 10, 2026 Session details

AI Agents Face-Off: Same App, Multiple Frameworks

Elaine Dias Batista

Which mobile framework wins when AI writes all the code? Discover how this multi-agent showdown proves that generating functional apps requires deeply restrictive, deterministic development environments.

Pause
Mute Enter Fullscreen
#1 about 4 min

Escaping human bias when comparing mobile frameworks

Delegating identical mobile app builds to artificial agents avoids personal framework biases and yields measurable comparison data.

#2 about 5 min

Iterating from manual to automated code generation workflows

Evolving developer instructions and evaluation workflows from manual trials and basic text guidelines to formalized context files.

#3 about 3 min

Eliminating code evaluation bias and tracking execution costs

Evaluating earlier experiments reveals the anti-pattern of model self-grading and shifts measurement from token usage to distinct financial costs.

#4 about 4 min

Designing an automated system harness for continuous testing

Creating a multi-loop feedback architecture with external modules enables structured planning and fully automated interface testing validation.

#5 about 2 min

Judging code quality with anonymous language model tournaments

Establishing a ranking system where external models perform blind evaluations guarantees completely objective output code quality measurements.

#6 about 4 min

Troubleshooting rogue agent behavior and mobile automation challenges

Overcoming unexpected complications where autonomous software changed accessibility features, triggered lock screens, and spawned unauthorized emulators.

#7 about 3 min

Analyzing benchmark results for coding agents and frameworks

Performance metrics demonstrate how specific model variants create higher quality applications with fewer errors and overall lower generation costs.

#8 about 2 min

Balancing deterministic framework architecture and probabilistic code generation

Implementing strict standard boilerplate within target frameworks simplifies functionality expectations compared to probabilistic raw project creation.

#9 about 3 min

Core engineering principles for an automated testing harness

Managing dynamic generation output requires strictly deterministic structures, environment version pinning, and disciplined context prompt budgeting.

#10 about 3 min

Assembling the complete execution toolchain for local agents

Executing identical software construction tests demands extensive orchestration scripts, target dependencies, specialized profilers, and platform build utilities.

Matching moments

3:06 min

The role of artificial intelligence in mobile development

Sasha Denisov Sasha Denisov

2:34 min

Balancing developer autonomy with the adoption of coding agents

Clemens Wasner Clemens Wasner +4 · World Congress 2026 Europe

2:06 min

Rethinking team structures around AI agent capabilities

Mike Mike · World Congress 2025

11:34 min

Rapid prototyping with AI agents at hackathons

Chris Heilmann +2 · LIVE

2:35 min

Evaluating frameworks and abstraction levels for building AI agents

Ahmad Adel Ahmad Adel · Europe 2026 Virtual

1:53 min

Validating autonomous code generation with robust automated testing

Ahmed Tikiwa Ahmed Tikiwa · World Congress 2026 Europe

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 24, 2026 · 13:30–14:00

Stage 1

How AI Agents Tripled Our Test Coverage on a 1.8M-Line iOS Codebase

Kush Agrawal

Staff Software Engineer, Platform

Kush Agrawal
Open session

World Congress 2026 North America

September 24, 2026 · 11:00–13:00

Stage 11

RoboCoders: Judgment Day: AI-Assisted Engineering Applied - The Battle of Agents

Baruch Sadogursky, Viktor Gamov

Baruch Sadogursky
Viktor Gamov
Open session

World Congress 2026 North America

September 25, 2026 · 13:30–14:00

Stage 7

Agents Can't Iterate Against Tests That Lie

Rocky Warren

Senior Staff Software Engineer at Clipboard

Rocky Warren
Open session

World Congress 2026 North America

September 24, 2026 · 11:40–12:10

Stage 4

Taming Rogue Agents: Observability-Driven Evaluation for Production Reliability

Anagha Rumade, Anjana Umapathy, Apoorva Jaiswal

Anagha Rumade
Anjana Umapathy
Apoorva Jaiswal
Open session

World Congress 2026 North America

September 24, 2026 · 11:20–11:25

Outdoor Stage

Finding the Edges: Testing, Evaluating, and Monitoring Voice AI Agents Before Your Users Do

Matt Wyman

CEO at Okareo

Matt Wyman
Open session

World Congress 2026 North America

September 25, 2026 · 15:30–16:00

Stage 5

21 Experiments in Six Weeks: A Playbook for Improving Your AI Agent

Sofia Rest

Software Engineer at Sentry

Sofia Rest