World Congress 2026 Europe • Jul 10, 2026 • Session details

Hack Me If You Can: Designing Unbreakable LLM Guardrails

Cansu Kavili Örnek

System prompts will not stop prompt injection. Stop wasting costly GPU cycles on malicious requests. Build unbreakable GenAI guardrails using a layered defense architecture.

Pause
Mute Enter Fullscreen
#1 about 3 min

Why system prompts fail against LLM jailbreak attempts

Persistent threat actors can easily bypass basic system instructions to exploit generative language models.

#2 about 4 min

Security and cost implications of unguarded language models

Unrestricted models process out-of-domain requests that leak sensitive data and waste expensive compute resources.

#3 about 2 min

Architectural overview of the guardrails intercept pattern

Intercepting model traffic with a specialized orchestration layer prevents unauthorized interactions without altering the underlying model.

#4 about 6 min

Interactive demonstration of intercepting unauthorized model prompts

A targeted demonstration shows how conditional rules block out-of-domain questions and language-switching bypass techniques.

#5 about 4 min

Implementing regex, classifiers, and LLM-as-a-judge guardrails

Teams construct layered defenses by mixing cheap regex rules with lightweight structural classifiers and analytical evaluation models.

#6 about 5 min

Deploying open-source guardrails and continuous evaluation tools

Integrating orchestration tools and testing hubs secures chat completion endpoints against continuous red teaming attacks.

#7 about 3 min

Scaling modular AI guardrails across multiple enterprise teams

Centralized platform teams configure baseline security protocols while authorizing product groups to deploy domain-specific compliance rules.

#8 about 3 min

Tracking the enterprise maturity journey for AI defenses

Maturing AI operations involve transitioning from raw endpoints to an optimized service that balances latency constraints and application safety.

#9 about 5 min

Handling data access, custom classifiers, and continuous evaluation

Training lightweight custom classifiers on synthetic data provides affordable security checks before scaling into robust evaluation benchmarks.

Matching moments

2:15 min

Establishing guardrails and infrastructure for AI models

Alexandre Guenoun Alexandre Guenoun +3 · World Congress 2026 Europe

2:43 min

Setting effective guardrails for enterprise agentic AI adoption

Julia Kordick Julia Kordick · Coffee With Developers

2:13 min

Identifying and hardening against generative AI risks

Rebekka Weiss Rebekka Weiss +1 · World Congress 2025

1:20 min

Utilizing industry threat models for AI security

Balázs Kiss · World Congress 2023

4:15 min

Security integration and AI skepticism in developer tooling

Chris Heilmann Chris Heilmann +2 · LIVE

4:56 min

Q&A on platform discovery, security, and prompt injection

Rishabh Budhiraja Rishabh Budhiraja +1 · World Congress 2026 Europe