adversarial-rl

AI Audit & Security: From Reinforcement Learning to LLM Agents

This domain hosted the IJCAI 2022 tutorial on adversarial reinforcement learning. It is now an independent reference on how AI systems are evaluated, audited, and defended.

See the coverage map

Live demo: evasion attack on a malware classifier

The input is almost unchanged. The verdict flips. This is an adversarial example.

From the 2022 tutorial to the full field

The IJCAI 2022 tutorial had five parts. Each part now has a matching page on this site.

Tutorial partTopicPage on this site
Part 0Introduction and preliminariesAI Audit & Security/en/ai-security-audit/
Part 1Test-time attacks and defensesAdversarial attacks on reinforcement learning/en/adversarial-attacks-reinforcement-learning/
Part 2Training-time attacksData poisoning and backdoor attacks/en/data-poisoning-backdoor-attacks/
Part 3Training-time defensesAI guardrails and monitoring/en/ai-guardrails-monitoring/
Part 4Adversarial multi-agent RLAI agent security audit/en/ai-agent-security-audit/

Three older domains lead here

Each domain worked on a part of the same question: how far can a system be trusted when someone tries to break it.

Former domainOriginal subjectWhy it belongs hereLanding page
aofa2007.orgAnalysis of algorithmsWorst-case analysis defines robustness: what an attacker can force, not what usually happens.Analysis of algorithms and worst-case robustness/en/analysis-of-algorithms-worst-case-robustness/
strategicreasoning.netStrategic reasoning, agents, game theoryAttackers and defenders are strategic agents. Game theory models both.Strategic reasoning and multi-agent security/en/strategic-reasoning-multi-agent-security/
wasaconf.orgAI penetration testing, RAG evaluationThe practical side: testing a deployed AI system and reporting the findings.AI penetration testing and audit/en/ai-penetration-testing-audit/