Adversarial Attacks on Reinforcement Learning

How can an attacker fool an AI system at test time, during inference? By perturbing the input—slightly. Evasion attacks find the boundary between safe decisions and failure. This is Part 1 of the IJCAI 2022 tutorial.

See the workflow →

The Four Phases

An attack follows a repeatable, documented process.

  • Attack design
  • Perturbation generation
  • Success evaluation
  • Robustness metrics

Core Concepts

The foundations of understanding adversarial threats.

Why Test-Time Attacks Matter

A deployed RL agent makes decisions based on observations. An attacker can perturb those observations—slightly and often imperceptibly—to change the agent's decision. The goal: find the boundary between safe and unsafe input space.

  • Find decision boundaries in high-dimensional input space
  • Measure robustness to adversarial perturbations
  • Evaluate worst-case performance under attack
  • Design robust policies that resist evasion

Attack Methodology

An adversarial attack has four phases: design (choose attack strategy), perturb (generate adversarial examples), evaluate (measure success rate), and verify (ensure robustness). Each attack tests a specific vulnerability.

  • FGSM and PGD: gradient-based perturbations
  • Certified defenses: provable robustness bounds
  • Evaluation metrics: success rate, perturbation magnitude
  • Robustness verification: formal methods and testing

The Attack Workflow

Each phase is methodical and repeatable. No phase is skipped; each has clear objectives.

Design attack

Choose the threat model: what can the attacker observe and modify?

Generate perturbations

Compute adversarial examples using gradient-based or optimization methods.

Evaluate success

Measure attack success rate and average perturbation magnitude.

Verify robustness

Test whether defenses actually prevent similar attacks.

Related Pages

Deep dives into complementary attack and defense topics.

AI Audit & Security Framework

The methodology that guides adversarial attack campaigns.

Data Poisoning and Backdoors

Training-time attacks on the same systems.

AI Guardrails and Monitoring

Defense mechanisms for runtime protection.

OWASP LLM Top 10

LLM-specific evasion attacks.