World Congress 2026 Europe - Virtual Stage Jun 30, 2026 Session details

Load Testing AI: Aiming at a Moving Target

Heather Thacker

Are your traditional load tests hiding massive AI retry storms? Learn why LLM architectures require entirely new performance metrics to prevent violent latency cascades and runaway costs.

Pause
Mute Enter Fullscreen
#1 about 2 min

Why traditional API load testing fails for AI endpoints

Applying static web request metrics to AI systems obscures critical latency and cost issues in production.

#2 about 5 min

Mechanical differences between web APIs and AI traffic

AI endpoints introduce variable workload distributions, bounded GPU concurrency, and token-dependent generation costs.

#3 about 5 min

Five unique failure modes in production AI systems

Hidden queuing delays, token truncation, and retry storms introduce cascading performance degradation under heavy load.

#4 about 6 min

Modeling realistic AI workloads and user behavior testing

Using randomized prompt datasets and open workload models accurately simulates inference queuing and cache hit variability.

#5 about 3 min

Evaluating latency, cost, and resilience beyond request throughput

Tracking time to first token alongside generation limits transforms unpredictable billing into measurable performance guardrails.

#6 about 8 min

Contrasting static versus randomized prompt load testing outcomes

Replacing a single fixed string with a diverse prompt corpus reveals true tail latencies and processing times.

#7 about 6 min

Simulating AI system cascades and graceful load shedding

Pushing an endpoint beyond its inference capacity verifies whether bounded queues and circuit breakers prevent complete outages.

#8 about 5 min

Establishing AI-native service level objectives and safety guardrails

Implementing token-based performance targets and continuous cost observability ensures long-term operational resilience for AI services.

Matching moments

2:13 min

Identifying and hardening against generative AI risks

Rebekka Weiss Rebekka Weiss +1 · World Congress 2025

3:46 min

Integrating AI into web performance engineering workflows

Perf + AI

2:42 min

Managing AI token costs through strategic test design

Jakub Janczyk Jakub Janczyk · World Congress 2026 Europe

4:12 min

Measuring generative AI impact on team productivity

Justin Reock Justin Reock · World Congress 2026 Europe

54 sec

Architecting generative AI systems by targeted user load

Stan Girard Stan Girard · World Congress 2024

1:38 min

Scaling bottlenecks in generative AI applications

Stan Girard Stan Girard · World Congress 2024

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 25, 2026 · 11:40–12:10

Stage 2

Reinventing Testing Practices in the AI Era

Eric Deandrea

Java Champion & Senior Principal Software Engineer at IBM

Eric Deandrea
Open session

World Congress 2026 North America

September 24, 2026 · 14:10–14:40

Stage 1

Anatomy of an AI Request: Where Latency and Cost Are Really Born

Dan Fu

VP of Kernels at Together AI

Dan Fu
Open session

World Congress 2026 North America

September 24, 2026 · 16:50–17:20

Stage 6

Who Tests the AI? Building Trustworthy AI Systems at Enterprise Scale

Him Raj Singh

Manager, Software Engineer at PayPal

Him Raj Singh
Open session

World Congress 2026 North America

September 23, 2026 · 10:45–12:45

Stage 10

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

September 24, 2026 · 14:10–14:40

Stage 7

Designing High-Performance AI APIs: Lessons from Serving Millions of Real-Time Requests

Wayne Liu

Chief Growth Officer and Americas President of Perfect Corp.

Wayne Liu
Open session

World Congress 2026 North America

September 25, 2026 · 11:40–12:10

Tech Leaders Stage

The AI Proving Ground

Alexius Wronka

CTO of Data and Growth at Invisible Technologies

Alexius Wronka