AI Audit & Security: From Reinforcement Learning to LLM Agents
This domain hosted the IJCAI 2022 tutorial on adversarial reinforcement learning. It is now an independent reference on how AI systems are evaluated, audited, and defended.
See the coverage mapLive demo: evasion attack on a malware classifier
The input is almost unchanged. The verdict flips. This is an adversarial example.
From the 2022 tutorial to the full field
The IJCAI 2022 tutorial had five parts. Each part now has a matching page on this site.
| Tutorial part | Topic | Page on this site |
|---|---|---|
| Part 0 | Introduction and preliminaries | AI Audit & Security/en/ai-security-audit/ |
| Part 1 | Test-time attacks and defenses | Adversarial attacks on reinforcement learning/en/adversarial-attacks-reinforcement-learning/ |
| Part 2 | Training-time attacks | Data poisoning and backdoor attacks/en/data-poisoning-backdoor-attacks/ |
| Part 3 | Training-time defenses | AI guardrails and monitoring/en/ai-guardrails-monitoring/ |
| Part 4 | Adversarial multi-agent RL | AI agent security audit/en/ai-agent-security-audit/ |
Three older domains lead here
Each domain worked on a part of the same question: how far can a system be trusted when someone tries to break it.
| Former domain | Original subject | Why it belongs here | Landing page |
|---|---|---|---|
| aofa2007.org | Analysis of algorithms | Worst-case analysis defines robustness: what an attacker can force, not what usually happens. | Analysis of algorithms and worst-case robustness/en/analysis-of-algorithms-worst-case-robustness/ |
| strategicreasoning.net | Strategic reasoning, agents, game theory | Attackers and defenders are strategic agents. Game theory models both. | Strategic reasoning and multi-agent security/en/strategic-reasoning-multi-agent-security/ |
| wasaconf.org | AI penetration testing, RAG evaluation | The practical side: testing a deployed AI system and reporting the findings. | AI penetration testing and audit/en/ai-penetration-testing-audit/ |
Every field of AI security, one page each
Twelve pages cover attacks, audit method, defenses and frameworks.
AI Audit & Security FrameworkScope, method, evidence and report of an audit./en/ai-security-audit/LLM Red TeamingStructured attack campaigns against language models./en/llm-red-teaming/Prompt Injection AttacksDirect, indirect and RAG injection, with tests and fixes./en/prompt-injection-attacks/AI Agent Security AuditTool use, privileges, memory poisoning, orchestration./en/ai-agent-security-audit/Adversarial Attacks on RLTest-time and training-time attacks on policies./en/adversarial-attacks-reinforcement-learning/Data Poisoning and BackdoorsCorrupted training data, trojans, sleeper behavior./en/data-poisoning-backdoor-attacks/Model Extraction & Privacy AttacksModel theft, membership inference, data leakage./en/model-extraction-privacy-attacks/AI Supply Chain SecurityDatasets, pretrained models, dependencies, checkpoints./en/ai-supply-chain-security/AI Guardrails and MonitoringFiltering, drift detection, runtime alerts, robust training./en/ai-guardrails-monitoring/OWASP LLM Top 10The ten LLM risks, mapped to audit tests./en/owasp-llm-top-10/MITRE ATLAS FrameworkAdversary tactics and techniques against ML systems./en/mitre-atlas-framework/AI Governance & EU AI ActObligations, controls and audit evidence./en/ai-governance-eu-ai-act-iso-42001/