World Congress 2025 Aug 20, 2025 Session details

You are not my model anymore - understanding LLM model behavior

Andreas Erben

A silent filter update once caused Microsoft Copilot to completely hide internal documents. Learn how automated evaluations and uncensored testing models protect applications from unpredictable LLM shifts.

Pause
Mute Enter Fullscreen
#1 about 3 min

Silent filter updates causing unexpected file hiding

Microsoft's undocumented filter implementations for jailbreaks can silently hide entire documents from searches.

#2 about 4 min

Generating tokens and shaping model behaviors

Large language models generate outputs as token probability distributions and require structured alignment to behave predictably.

#3 about 2 min

Navigating hidden variables in closed source model architectures

Major artificial intelligence labs keep their model topologies and runtime configurations hidden behind layers of application code.

#4 about 3 min

Navigating undocumented content filters in cloud infrastructures

Cloud providers enforce restrictive content filters that silently block prompts and inadvertently enable second-order data-hiding attacks.

#5 about 3 min

Managing rapid lifecycles and shifting model behaviors

Frequent model deprecations and automated version upgrades result in unannounced changes to expected software behavior.

#6 about 3 min

Understanding the fragile alignment of language models

Aligned behavior represents a thin veneer over unpredictable patterns that can be bypassed using grammar-based jailbreaks.

#7 about 2 min

Mapping internal model concepts to activation weights

Observing internal activation weights during text generation facilitates the early detection of hallucinations and distinct model personalities.

#8 about 4 min

Building evaluation frameworks for automated regression testing

Continuous regression testing using automated prompt permutations prevents systemic failures when expected model versions upgrade rapidly.

#9 about 5 min

Optimizing prompts and testing against unannounced filter updates

Utilizing uncensored evaluation models and explicit zero-shot structures helps developers test consistently against hidden cloud filter updates.

Matching moments

5:25 min

Addressing core challenges in large language model deployments

Vijay Krishan Gupta +1 · LIVE

2:52 min

Mitigating indirect prompt injection in language models

Jose Manuel Ortega Jose Manuel Ortega · Europe 2026 Virtual

4:29 min

Designing AI applications defensively for inevitable failures

Krzysztof Cieślak Krzysztof Cieślak · World Congress 2026 Europe

4:25 min

Overcoming AI hallucinations and restrictive content guardrails

Perf + AI

2:18 min

Exploiting exposed language models in commercial applications

Liran Tal Liran Tal · LIVE

1:40 min

Evaluating ongoing security risks in large language models

Sebastian Schrittwieser · World Congress 2023

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 24, 2026 · 17:30–18:00

Stage 6

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

September 24, 2026 · 13:30–14:00

Stage 9

Understanding LLM Architectures: Inside the Design of Modern Models

Jofia Jose Prakash

Director - AI & Governance at Humanity + AI, Inc

Jofia Jose Prakash
Open session

World Congress 2026 North America

September 25, 2026 · 11:40–12:10

Stage 9

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

September 24, 2026 · 16:50–17:20

Stage 5

Context Engineering Kung Fu

Carl Lapierre

Tech Lead and AI Engineer at Osedea

Carl Lapierre
Open session

World Congress 2026 North America

September 25, 2026 · 09:00–09:30

Stage 6

Red Teaming Your LLM App -- A Hands-On Threat Model You Can Reuse

Saloni Garg

Senior ML Engineer at Adobe

Saloni Garg
Open session

World Congress 2026 North America

September 24, 2026 · 14:25–14:35

Outdoor Stage

SecurePrompt: Building a Pre-Flight Security Layer for Agentic AI

Ravi Sastry Kadali

AI/ML Engineer at General Motors

Ravi Sastry Kadali