LLM Red Teaming
Red teaming a language model means running structured attack campaigns to find failure modes. Jailbreaks, prompt injection, role-play scenarios—each tests the boundaries of model behavior and safety.
See the workflow →The Four Phases
An attack follows a repeatable, documented process.
- Threat modeling
- Attack design
- Execution
- Impact assessment
Core Concepts
The foundations of understanding these vulnerabilities.
Why Red Teaming Matters
Language models are not aligned by default. They can be tricked into harmful outputs through prompt manipulation, role-play, or adversarial jailbreaks. Red teaming finds these vulnerabilities before deployment.
- Discover jailbreak methods and workarounds
- Identify topics where model guardrails fail
- Measure robustness to adversarial prompts
- Create evidence for safety improvements
Red Team Methodology
A red team follows a structured process: model (understand capabilities), attack (design prompts to bypass guardrails), measure (quantify success), and document (record findings). Each attack tests a specific vulnerability.
- Direct jailbreaks: explicit requests for harmful content
- Indirect jailbreaks: role-play, hypotheticals, framing
- Prompt injection: hiding attacks in benign-looking input
- Behavioral testing: edge cases, unusual inputs, stress tests
The Attack Workflow
Each phase is methodical and repeatable. No phase is skipped; each has clear objectives.
Model capability scan
What can the model do? What guardrails exist?
Attack design
Craft prompts that bypass safety mechanisms.
Execute campaigns
Run structured jailbreak attempts and log results.
Assess impact
Measure success rate, severity, and reproducibility.
Related Pages
Deep dives into related attack and defense topics.
Prompt Injection Attacks
How to embed attacks in multi-step conversations.
AI Audit & Security Framework
The methodology that guides red team campaigns.
LLM-Specific Vulnerabilities
Model-specific failure modes and workarounds.
OWASP LLM Top 10
The ten LLM risks, prioritized by red team findings.