World Congress 2023 Sep 27, 2023

ChatGPT, ignore the above instructions! Prompt injection attacks and how to avoid them.

Sebastian Schrittwieser

Are your AI system instructions leaking as public knowledge? Traditional filtering fails against prompt injections, but dual LLM architectures can isolate untrusted data and secure your application.

Pause
Mute Enter Fullscreen
#1 about 3 min

Growth of large language models and security awareness

The rapid adoption of artificial intelligence introduces significant security challenges that developers often overlook initially.

#2 about 2 min

Executing basic prompt injection attacks through translation tasks

Conflicting instructions in user input can successfully override a base language model's original operational directive.

#3 about 2 min

Understanding context and system prompts in applications

Defining boundaries between underlying developer instructions and incoming user input helps identify the source of intended model behavior.

#4 about 2 min

Real-world prompt injection attacks on social media bots

Real-world incidents demonstrate how targeted text inputs can easily manipulate an automated bot into generating non-relevant text.

#5 about 3 min

Extracting developer instructions via context leak attacks

Malicious users can force a model to expose its underlying instructions and sensitive logic by exploiting specific query structures.

#6 about 2 min

The inability to secure sensitive information in system prompts

Developers must treat all contextual instructions and system prompts as publicly accessible data since evasion techniques remain unpredictable.

#7 about 4 min

Exploiting external plugins with untrusted input manipulation

Connecting language models to external data sources creates dangerous new avenues for automated injection attacks.

#8 about 3 min

Comparing prompt injection to traditional SQL injection attacks

The fundamental lack of separation between developer code and user input in language models mirrors the mechanics of SQL injection.

#9 about 3 min

Falling short with traditional input and output filtering

Standard security techniques like blacklisting fail against the endless linguistic complexities and encoding methods of natural language prompts.

#10 about 2 min

Assessing the limitations of prompt reordering and confirmation dialogues

Placing user input before the system context or adding strict confirmation dialogues yields inadequate security against determined attackers.

#11 about 4 min

Isolating untrusted data using the dual language model architecture

A structural approach using privileged and quarantined models safely processes potentially malicious external data by enforcing strict separation.

#12 about 2 min

Evaluating ongoing security risks in large language models

Large language models remain highly vulnerable to untrusted input manipulations that override intended application behavior.

#13 about 4 min

Answering questions on prompt intent and data tagging

Exploring alternative defense mechanisms like XML tagging and intent analysis reveals the difficulty of detecting malicious instructions automatically.

Matching moments

2:46 min

Bypassing language model safeguards utilizing contextual prompt injection attacks

Chris Heilmann +1 · LIVE

5:44 min

Prompt injection vulnerabilities and contextual mitigation testing challenges

Mirko Ross · WWC 2023

3:04 min

Prompt injections bypassing language model service guardrails

Ramona Schwering Ramona Schwering · WWC Europe 2026

2:20 min

Why traditional security fails against large language models

Péter Farkas Péter Farkas · Europe 2026 Virtual

3:29 min

Identifying and defining prompt injection execution vulnerabilities

Mackenzie Mackenzie · WWC 2024

4:19 min

Prompt engineering techniques and security vulnerabilities

Aarno Aukia · LIVE

Upcoming sessions on this topic

Open session

World Congress 2026 North America

SecurePrompt: Building a Pre-Flight Security Layer for Agentic AI

Ravi Sastry Kadali

AI/ML Engineer at General Motors

Ravi Sastry Kadali
Open session

World Congress 2026 North America

The Things Your AI Isn't Telling You

Desmond Lamptey

Lead Software Engineer @ Capital One

Desmond Lamptey
Open session

World Congress 2026 North America

Red Teaming Your LLM App -- A Hands-On Threat Model You Can Reuse

Saloni Garg

Senior ML Engineer at Adobe

Saloni Garg
Open session

World Congress 2026 North America

Context Engineering Kung Fu

Carl Lapierre

Tech Lead and AI Engineer at Osedea

Carl Lapierre
Open session

World Congress 2026 North America

Responsible AI Architecture with Zero Trust Agents

Ashok Prakash

Staff ML Engineer at Apple

Ashok Prakash
Open session

World Congress 2026 North America

Know Your Enemies: Live Exploit of a PHP Engine Security Breach

Alexandre Daubois

CTO of Les-Tilleuls.coop / Symfony Core Team / PHP & FrankenPHP Core Maintainer

Alexandre Daubois