Model Extraction & Privacy Attacks

An attacker can steal a model by querying it repeatedly and replicating its behavior. Or they can infer membership in the training set. Both are privacy violations.

See the workflow →

The Four Phases

An attack follows a repeatable, documented process.

  • Target reconnaissance
  • Query strategy
  • Model reconstruction
  • Exploitation

Core Concepts

The foundations of understanding these vulnerabilities.

Why Model Extraction Matters

Model weights are IP. If an API exposes the model's predictions, an attacker can steal the entire model through repeated queries. Additionally, training data privacy can be violated through membership inference attacks.

  • Steal proprietary model weights through API access
  • Infer whether specific samples were in the training set
  • Extract sensitive patterns or decision rules
  • Violate privacy of training data contributors

Extraction & Privacy Methodology

Model theft has four phases: reconnaissance (profile the model), query (collect predictions), distill (train a substitute model), and exploit (use the stolen model). Privacy attacks work similarly.

  • Extraction: FGSM queries, adversarial examples, distillation
  • Membership inference: differential privacy violations
  • Training data reconstruction: GAN-based inversion attacks
  • Privacy defenses: differential privacy, output perturbation

The Attack Workflow

Each phase is methodical and repeatable. No phase is skipped; each has clear objectives.

Reconnaissance

What is the model architecture? How does it respond to queries?

Collect predictions

Query the API systematically to build a dataset of input-output pairs.

Train substitute

Use collected data to train a replica model that mimics the target.

Exploit or analyze

Use the stolen model for attacks or perform membership inference.

Related Pages

Deep dives into related attack and defense topics.

AI Supply Chain Security

Protecting model checkpoints and preventing unauthorized access.

AI Audit & Security Framework

Testing for model extraction vulnerabilities.

Data Poisoning and Backdoors

Defense against model theft through robust training.

AI Governance & EU AI Act

Privacy regulations and model ownership rights.