Prompt Injection Attacks
Prompt injection embeds attack commands in user input, system prompts, or retrieved context. An attacker tricks the model into following injected instructions instead of the original task.
See the workflow →The Four Phases
An attack follows a repeatable, documented process.
- Injection vector identification
- Payload design
- Execution
- Containment
Core Concepts
The foundations of understanding these vulnerabilities.
Why Prompt Injection Matters
A model follows instructions. If those instructions come from untrusted sources—user input, API responses, database records—an attacker can override the original task. Prompt injection is the most reliable LLM attack.
- Redirect model behavior through malicious input
- Bypass safety guidelines and content policies
- Extract sensitive information via prompt manipulation
- Chain attacks: injection → data exfiltration → code execution
Injection Attack Methodology
Prompt injection has three forms: direct (user supplies malicious prompt), indirect (attacker controls a data source), and RAG injection (attack hidden in retrieved context). Each requires different detection and mitigation.
- Direct injection: prompt override, jailbreak markers
- Indirect injection: poisoned API responses, database records
- RAG injection: attacks in retrieved documents or search results
- Chaining: multiple injections across tool calls
The Attack Workflow
Each phase is methodical and repeatable. No phase is skipped; each has clear objectives.
Map injection vectors
What data sources does the model trust? Where is input unsanitized?
Design payloads
Craft injection commands that override the original task.
Execute attacks
Test injections with different payloads and contexts.
Implement mitigations
Add input validation, prompt disclaimers, sandboxing.
Related Pages
Deep dives into related attack and defense topics.
LLM Red Teaming
Structured campaigns to find injection vulnerabilities.
AI Agent Security Audit
Tool-using agents vulnerable to prompt injection.
AI Audit & Security Framework
Testing methodology for injection robustness.
OWASP LLM Top 10
Prompt injection as OWASP 04 - Insecure Output Handling.