World Congress 2026 Europe Jul 10, 2026 Session details

Hack Me If You Can: Designing Unbreakable LLM Guardrails

Cansu Kavili Örnek

System prompts will not stop prompt injection. Stop wasting costly GPU cycles on malicious requests. Build unbreakable GenAI guardrails using a layered defense architecture.

Pause
Mute Enter Fullscreen
#1 about 3 min

Why system prompts fail against LLM jailbreak attempts

Persistent threat actors can easily bypass basic system instructions to exploit generative language models.

#2 about 4 min

Security and cost implications of unguarded language models

Unrestricted models process out-of-domain requests that leak sensitive data and waste expensive compute resources.

#3 about 2 min

Architectural overview of the guardrails intercept pattern

Intercepting model traffic with a specialized orchestration layer prevents unauthorized interactions without altering the underlying model.

#4 about 6 min

Interactive demonstration of intercepting unauthorized model prompts

A targeted demonstration shows how conditional rules block out-of-domain questions and language-switching bypass techniques.

#5 about 4 min

Implementing regex, classifiers, and LLM-as-a-judge guardrails

Teams construct layered defenses by mixing cheap regex rules with lightweight structural classifiers and analytical evaluation models.

#6 about 5 min

Deploying open-source guardrails and continuous evaluation tools

Integrating orchestration tools and testing hubs secures chat completion endpoints against continuous red teaming attacks.

#7 about 3 min

Scaling modular AI guardrails across multiple enterprise teams

Centralized platform teams configure baseline security protocols while authorizing product groups to deploy domain-specific compliance rules.

#8 about 3 min

Tracking the enterprise maturity journey for AI defenses

Maturing AI operations involve transitioning from raw endpoints to an optimized service that balances latency constraints and application safety.

#9 about 5 min

Handling data access, custom classifiers, and continuous evaluation

Training lightweight custom classifiers on synthetic data provides affordable security checks before scaling into robust evaluation benchmarks.

Matching moments

2:15 min

Establishing guardrails and infrastructure for AI models

Alexandre Guenoun Alexandre Guenoun +3 · WWC Europe 2026

2:43 min

Setting effective guardrails for enterprise agentic AI adoption

Julia Kordick Julia Kordick · Coffee With Developers

2:13 min

Identifying and hardening against generative AI risks

Rebekka Weiss Rebekka Weiss +1 · WWC 2025

1:20 min

Utilizing industry threat models for AI security

Balázs Kiss · WWC 2023

4:15 min

Security integration and AI skepticism in developer tooling

Chris Heilmann +2 · LIVE

4:56 min

Q&A on platform discovery, security, and prompt injection

Rishabh Budhiraja Rishabh Budhiraja +1 · WWC Europe 2026

Upcoming sessions on this topic

Open session

World Congress 2026 North America

SecurePrompt: Building a Pre-Flight Security Layer for Agentic AI

Ravi Sastry Kadali

AI/ML Engineer at General Motors

Ravi Sastry Kadali
Open session

World Congress 2026 North America

Red Teaming Your LLM App -- A Hands-On Threat Model You Can Reuse

Saloni Garg

Senior ML Engineer at Adobe

Saloni Garg
Open session

World Congress 2026 North America

The Things Your AI Isn't Telling You

Desmond Lamptey

Lead Software Engineer @ Capital One

Desmond Lamptey
Open session

World Congress 2026 North America

Securing AI Agent Infrastructure: Identity, Attestation, and Trust at Scale

Abdel Fane

Founder of OpenA2A

Abdel Fane
Open session

World Congress 2026 North America

Who Tests the AI? Building Trustworthy AI Systems at Enterprise Scale

Him Raj Singh

PayPal, Manager, Software Engineer

Him Raj Singh
Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong