AI Guardrails and Monitoring
How do you defend a model against attacks? Guardrails are protective mechanisms built into training. Monitoring detects anomalies at runtime. Together, they create a defense-in-depth strategy. This is Part 3 of the IJCAI 2022 tutorial.
See the workflow →The Four Phases
A defense strategy follows a repeatable, documented process.
- Defense design
- Robust training
- Monitoring setup
- Incident response
Core Concepts
The foundations of building secure, robust AI systems.
Why Defenses Matter
A defended model resists attacks by design. Defenses include robust training (learning from adversarial examples), certified robustness (provable bounds), and runtime monitoring (detecting anomalies). No defense is perfect, but a layered approach raises the cost of attack.
- Reduce success rate of known attacks
- Build measurable robustness guarantees
- Detect anomalies and trigger alerts
- Maintain security posture over time
Defense Strategy
Defense has four phases: design (choose defense mechanisms), train (implement robust training), monitor (set up runtime detection), and respond (react to alerts). Each phase reduces attack surface and response time.
- Input filtering: validation, normalization, obfuscation
- Robust training: adversarial examples, certified defenses
- Monitoring: input anomaly detection, output drift detection
- Incident response: logging, rollback, retraining
The Defense Workflow
Each phase is methodical and repeatable. No phase is skipped; each has clear objectives.
Design defense
What threats are we defending against? What trade-offs are acceptable?
Implement robust training
Train with adversarial examples or certified robustness methods.
Deploy monitoring
Set thresholds for input anomalies and output drift.
Respond to alerts
Log incidents, trigger escalation, plan retraining.
Related Pages
Deep dives into attack and defense contexts.
AI Audit & Security Framework
How to test whether defenses actually work.
Adversarial Attacks on RL
The test-time threats that defenses must withstand.
Data Poisoning and Backdoors
Training-time threats that defenses must prevent.
OWASP LLM Top 10
LLM-specific defense strategies.