AI Guardrails and Monitoring

How do you defend a model against attacks? Guardrails are protective mechanisms built into training. Monitoring detects anomalies at runtime. Together, they create a defense-in-depth strategy. This is Part 3 of the IJCAI 2022 tutorial.

See the workflow →

The Four Phases

A defense strategy follows a repeatable, documented process.

  • Defense design
  • Robust training
  • Monitoring setup
  • Incident response

Core Concepts

The foundations of building secure, robust AI systems.

Why Defenses Matter

A defended model resists attacks by design. Defenses include robust training (learning from adversarial examples), certified robustness (provable bounds), and runtime monitoring (detecting anomalies). No defense is perfect, but a layered approach raises the cost of attack.

  • Reduce success rate of known attacks
  • Build measurable robustness guarantees
  • Detect anomalies and trigger alerts
  • Maintain security posture over time

Defense Strategy

Defense has four phases: design (choose defense mechanisms), train (implement robust training), monitor (set up runtime detection), and respond (react to alerts). Each phase reduces attack surface and response time.

  • Input filtering: validation, normalization, obfuscation
  • Robust training: adversarial examples, certified defenses
  • Monitoring: input anomaly detection, output drift detection
  • Incident response: logging, rollback, retraining

The Defense Workflow

Each phase is methodical and repeatable. No phase is skipped; each has clear objectives.

Design defense

What threats are we defending against? What trade-offs are acceptable?

Implement robust training

Train with adversarial examples or certified robustness methods.

Deploy monitoring

Set thresholds for input anomalies and output drift.

Respond to alerts

Log incidents, trigger escalation, plan retraining.

Related Pages

Deep dives into attack and defense contexts.

AI Audit & Security Framework

How to test whether defenses actually work.

Adversarial Attacks on RL

The test-time threats that defenses must withstand.

Data Poisoning and Backdoors

Training-time threats that defenses must prevent.

OWASP LLM Top 10

LLM-specific defense strategies.