Data Poisoning and Backdoor Attacks

An attacker can corrupt the training data itself—injecting poisons, trojans, and sleeper behaviors. The model learns from bad data and becomes unreliable. This is Part 2 of the IJCAI 2022 tutorial.

See the workflow →

The Four Phases

An attack follows a repeatable, documented process.

  • Data collection
  • Poison injection
  • Training & verification
  • Deployment

Core Concepts

The foundations of understanding adversarial threats.

Why Training-Time Attacks Matter

Training data is trust critical. If an attacker can poison the training set—either directly or through a supply chain vulnerability—they control what the model learns. Backdoors are hidden triggers that activate malicious behavior.

  • Corrupt model behavior without visible code changes
  • Hide attacks: poisoning affects a small fraction of data
  • Trigger backdoors on specific inputs (trojans)
  • Compromise entire supply chains (datasets, libraries)

Attack and Defense Methodology

Data poisoning has four phases: collection (identify data sources), injection (craft poisoned samples), training (monitor for backdoor), and deployment (detect triggers). Defenses must catch both overt and covert attacks.

  • Label flipping and feature poisoning attacks
  • Trojan triggers: physical triggers, semantic backdoors
  • Detection: anomaly detection in training, behavioral testing
  • Supply chain security: verify datasets, model checkpoints

The Attack Workflow

Each phase is methodical and repeatable. No phase is skipped; each has clear objectives.

Identify targets

Which datasets? Which models? What trust assumptions are broken?

Craft poisoned samples

Design backdoors: triggers, payloads, plausibility.

Inject and train

Add poison to training data; verify the backdoor works.

Deploy and trigger

Release the model; execute the backdoor on target inputs.

Related Pages

Deep dives into complementary attack and defense topics.

AI Audit & Security Framework

How to test for and detect poisoned models.

Adversarial Attacks on RL

Test-time evasion—the complement to training-time attacks.

AI Supply Chain Security

Defending against compromised datasets and models.

AI Guardrails and Monitoring

Runtime detection of backdoor activation.