LLM Red Teaming

Red teaming a language model means running structured attack campaigns to find failure modes. Jailbreaks, prompt injection, role-play scenarios—each tests the boundaries of model behavior and safety.

See the workflow →

The Four Phases

An attack follows a repeatable, documented process.

  • Threat modeling
  • Attack design
  • Execution
  • Impact assessment

Core Concepts

The foundations of understanding these vulnerabilities.

Why Red Teaming Matters

Language models are not aligned by default. They can be tricked into harmful outputs through prompt manipulation, role-play, or adversarial jailbreaks. Red teaming finds these vulnerabilities before deployment.

  • Discover jailbreak methods and workarounds
  • Identify topics where model guardrails fail
  • Measure robustness to adversarial prompts
  • Create evidence for safety improvements

Red Team Methodology

A red team follows a structured process: model (understand capabilities), attack (design prompts to bypass guardrails), measure (quantify success), and document (record findings). Each attack tests a specific vulnerability.

  • Direct jailbreaks: explicit requests for harmful content
  • Indirect jailbreaks: role-play, hypotheticals, framing
  • Prompt injection: hiding attacks in benign-looking input
  • Behavioral testing: edge cases, unusual inputs, stress tests

The Attack Workflow

Each phase is methodical and repeatable. No phase is skipped; each has clear objectives.

Model capability scan

What can the model do? What guardrails exist?

Attack design

Craft prompts that bypass safety mechanisms.

Execute campaigns

Run structured jailbreak attempts and log results.

Assess impact

Measure success rate, severity, and reproducibility.

Related Pages

Deep dives into related attack and defense topics.

Prompt Injection Attacks

How to embed attacks in multi-step conversations.

AI Audit & Security Framework

The methodology that guides red team campaigns.

LLM-Specific Vulnerabilities

Model-specific failure modes and workarounds.

OWASP LLM Top 10

The ten LLM risks, prioritized by red team findings.