World Congress 2026 Europe - Virtual Stage Jun 30, 2026 Session details

Load Testing AI: Aiming at a Moving Target

Heather Thacker

Are your traditional load tests hiding massive AI retry storms? Learn why LLM architectures require entirely new performance metrics to prevent violent latency cascades and runaway costs.

Pause
Mute Enter Fullscreen
#1 about 2 min

Why traditional API load testing fails for AI endpoints

Applying static web request metrics to AI systems obscures critical latency and cost issues in production.

#2 about 5 min

Mechanical differences between web APIs and AI traffic

AI endpoints introduce variable workload distributions, bounded GPU concurrency, and token-dependent generation costs.

#3 about 5 min

Five unique failure modes in production AI systems

Hidden queuing delays, token truncation, and retry storms introduce cascading performance degradation under heavy load.

#4 about 6 min

Modeling realistic AI workloads and user behavior testing

Using randomized prompt datasets and open workload models accurately simulates inference queuing and cache hit variability.

#5 about 3 min

Evaluating latency, cost, and resilience beyond request throughput

Tracking time to first token alongside generation limits transforms unpredictable billing into measurable performance guardrails.

#6 about 8 min

Contrasting static versus randomized prompt load testing outcomes

Replacing a single fixed string with a diverse prompt corpus reveals true tail latencies and processing times.

#7 about 6 min

Simulating AI system cascades and graceful load shedding

Pushing an endpoint beyond its inference capacity verifies whether bounded queues and circuit breakers prevent complete outages.

#8 about 5 min

Establishing AI-native service level objectives and safety guardrails

Implementing token-based performance targets and continuous cost observability ensures long-term operational resilience for AI services.

Matching moments

2:13 min

Identifying and hardening against generative AI risks

Rebekka Weiss Rebekka Weiss +1 · WWC 2025

3:46 min

Integrating AI into web performance engineering workflows

Perf + AI

2:42 min

Managing AI token costs through strategic test design

Jakub Janczyk Jakub Janczyk · WWC Europe 2026

4:12 min

Measuring generative AI impact on team productivity

Justin Reock Justin Reock · WWC Europe 2026

54 sec

Architecting generative AI systems by targeted user load

Stan Girard Stan Girard · WWC 2024

1:38 min

Scaling bottlenecks in generative AI applications

Stan Girard Stan Girard · WWC 2024

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Reinventing Testing Practices in the AI Era

Eric Deandrea

Java Champion & Senior Principal Software Engineer, IBM

Eric Deandrea
Open session

World Congress 2026 North America

Who Tests the AI? Building Trustworthy AI Systems at Enterprise Scale

Him Raj Singh

PayPal, Manager, Software Engineer

Him Raj Singh
Open session

World Congress 2026 North America

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

Designing APIs That Survive AI Agents at Scale

Phani Pendurthi

Mastercard, Principal Software Engineer

Phani Pendurthi
Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan