Data Poisoning and Backdoor Attacks
An attacker can corrupt the training data itself—injecting poisons, trojans, and sleeper behaviors. The model learns from bad data and becomes unreliable. This is Part 2 of the IJCAI 2022 tutorial.
See the workflow →The Four Phases
An attack follows a repeatable, documented process.
- Data collection
- Poison injection
- Training & verification
- Deployment
Core Concepts
The foundations of understanding adversarial threats.
Why Training-Time Attacks Matter
Training data is trust critical. If an attacker can poison the training set—either directly or through a supply chain vulnerability—they control what the model learns. Backdoors are hidden triggers that activate malicious behavior.
- Corrupt model behavior without visible code changes
- Hide attacks: poisoning affects a small fraction of data
- Trigger backdoors on specific inputs (trojans)
- Compromise entire supply chains (datasets, libraries)
Attack and Defense Methodology
Data poisoning has four phases: collection (identify data sources), injection (craft poisoned samples), training (monitor for backdoor), and deployment (detect triggers). Defenses must catch both overt and covert attacks.
- Label flipping and feature poisoning attacks
- Trojan triggers: physical triggers, semantic backdoors
- Detection: anomaly detection in training, behavioral testing
- Supply chain security: verify datasets, model checkpoints
The Attack Workflow
Each phase is methodical and repeatable. No phase is skipped; each has clear objectives.
Identify targets
Which datasets? Which models? What trust assumptions are broken?
Craft poisoned samples
Design backdoors: triggers, payloads, plausibility.
Inject and train
Add poison to training data; verify the backdoor works.
Deploy and trigger
Release the model; execute the backdoor on target inputs.
Related Pages
Deep dives into complementary attack and defense topics.
AI Audit & Security Framework
How to test for and detect poisoned models.
Adversarial Attacks on RL
Test-time evasion—the complement to training-time attacks.
AI Supply Chain Security
Defending against compromised datasets and models.
AI Guardrails and Monitoring
Runtime detection of backdoor activation.