World Congress 2025 Aug 20, 2025 Session details

You are not my model anymore - understanding LLM model behavior

Andreas Erben

A silent filter update once caused Microsoft Copilot to completely hide internal documents. Learn how automated evaluations and uncensored testing models protect applications from unpredictable LLM shifts.

Pause
Mute Enter Fullscreen
#1 about 3 min

Silent filter updates causing unexpected file hiding

Microsoft's undocumented filter implementations for jailbreaks can silently hide entire documents from searches.

#2 about 4 min

Generating tokens and shaping model behaviors

Large language models generate outputs as token probability distributions and require structured alignment to behave predictably.

#3 about 2 min

Navigating hidden variables in closed source model architectures

Major artificial intelligence labs keep their model topologies and runtime configurations hidden behind layers of application code.

#4 about 3 min

Navigating undocumented content filters in cloud infrastructures

Cloud providers enforce restrictive content filters that silently block prompts and inadvertently enable second-order data-hiding attacks.

#5 about 3 min

Managing rapid lifecycles and shifting model behaviors

Frequent model deprecations and automated version upgrades result in unannounced changes to expected software behavior.

#6 about 3 min

Understanding the fragile alignment of language models

Aligned behavior represents a thin veneer over unpredictable patterns that can be bypassed using grammar-based jailbreaks.

#7 about 2 min

Mapping internal model concepts to activation weights

Observing internal activation weights during text generation facilitates the early detection of hallucinations and distinct model personalities.

#8 about 4 min

Building evaluation frameworks for automated regression testing

Continuous regression testing using automated prompt permutations prevents systemic failures when expected model versions upgrade rapidly.

#9 about 5 min

Optimizing prompts and testing against unannounced filter updates

Utilizing uncensored evaluation models and explicit zero-shot structures helps developers test consistently against hidden cloud filter updates.

Matching moments

5:25 min

Addressing core challenges in large language model deployments

Vijay Krishan Gupta +1 · LIVE

4:29 min

Designing AI applications defensively for inevitable failures

Krzysztof Cieślak Krzysztof Cieślak · WWC Europe 2026

4:25 min

Overcoming AI hallucinations and restrictive content guardrails

Perf + AI

2:18 min

Exploiting exposed language models in commercial applications

Liran Tal Liran Tal · LIVE

1:40 min

Evaluating ongoing security risks in large language models

Sebastian Schrittwieser · WWC 2023

1:55 min

Identifying and fixing over-engineered AI calls through observability

diabhey diabhey · WWC 2025

Upcoming sessions on this topic

Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

Understanding LLM Architectures: Inside the Design of Modern Models

Jofia Jose Prakash

Enterprise AI Architect at American Chemical Society

Jofia Jose Prakash
Open session

World Congress 2026 North America

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

Context Engineering Kung Fu

Carl Lapierre

Tech Lead and AI Engineer at Osedea

Carl Lapierre
Open session

World Congress 2026 North America

Red Teaming Your LLM App -- A Hands-On Threat Model You Can Reuse

Saloni Garg

Senior ML Engineer at Adobe

Saloni Garg
Open session

World Congress 2026 North America

SecurePrompt: Building a Pre-Flight Security Layer for Agentic AI

Ravi Sastry Kadali

AI/ML Engineer at General Motors

Ravi Sastry Kadali