Robustness Leaderboard
Model × category.
Robustness is the fraction of attack traces the agent left not compromised, so higher is safer. The numbers below are placeholder fixture data, not a claimed benchmark; each cell opens that run in the Live Attack Replay.
Weakest0.29Model CASI06 · Memory & Context Poisoningnominal, greater than or equal to .80.50 to .79cautionbreach, less than .50Category names → ThreatsSOURCE · fixture
Swipe · 7 categories · model column stays pinned
| Model | ASI01 Agent Goal Hijack | ASI02 Tool Misuse and Exploitation | ASI03 Identity and Privilege Abuse | ASI04 Agentic Supply Chain Vulnerabilities | ASI05 Unexpected Code Execution (RCE) | ASI06 Memory & Context Poisoning | ASI10 Rogue Agents | Overall |
|---|---|---|---|---|---|---|---|---|
| Model A | 0.92 | 0.81 | 0.79 | 0.74 | 0.55 | 0.38 | 0.86 | 0.72 |
| Model B | 0.88 | 0.69 | 0.63 | 0.55 | 0.44 | 0.52 | 0.79 | 0.64 |
| Model C | 0.71 | 0.48 | 0.52 | 0.41 | 0.33 | 0.29 | 0.63 | 0.48 |
Cell → its run in Live Attack Replay