Adversarial Attacks on Reinforcement Learning
How can an attacker fool an AI system at test time, during inference? By perturbing the input—slightly. Evasion attacks find the boundary between safe decisions and failure. This is Part 1 of the IJCAI 2022 tutorial.
See the workflow →The Four Phases
An attack follows a repeatable, documented process.
- Attack design
- Perturbation generation
- Success evaluation
- Robustness metrics
Core Concepts
The foundations of understanding adversarial threats.
Why Test-Time Attacks Matter
A deployed RL agent makes decisions based on observations. An attacker can perturb those observations—slightly and often imperceptibly—to change the agent's decision. The goal: find the boundary between safe and unsafe input space.
- Find decision boundaries in high-dimensional input space
- Measure robustness to adversarial perturbations
- Evaluate worst-case performance under attack
- Design robust policies that resist evasion
Attack Methodology
An adversarial attack has four phases: design (choose attack strategy), perturb (generate adversarial examples), evaluate (measure success rate), and verify (ensure robustness). Each attack tests a specific vulnerability.
- FGSM and PGD: gradient-based perturbations
- Certified defenses: provable robustness bounds
- Evaluation metrics: success rate, perturbation magnitude
- Robustness verification: formal methods and testing
The Attack Workflow
Each phase is methodical and repeatable. No phase is skipped; each has clear objectives.
Design attack
Choose the threat model: what can the attacker observe and modify?
Generate perturbations
Compute adversarial examples using gradient-based or optimization methods.
Evaluate success
Measure attack success rate and average perturbation magnitude.
Verify robustness
Test whether defenses actually prevent similar attacks.
Related Pages
Deep dives into complementary attack and defense topics.
AI Audit & Security Framework
The methodology that guides adversarial attack campaigns.
Data Poisoning and Backdoors
Training-time attacks on the same systems.
AI Guardrails and Monitoring
Defense mechanisms for runtime protection.
OWASP LLM Top 10
LLM-specific evasion attacks.