Skip to content
MCPwn
SAMPLE

Red-team your MCP agent. Trust the verdict either way.

Point your own agent at an endpoint we host, and we serve it one of the OWASP Agentic Top 10 attacks and record every tool call it chooses to make.

Most tools only guess whether an agent was compromised. We tested our detector against attacks with known outcomes and measured how accurate it is, so you can trust its verdicts.

Measured finding Asked to fetch the invoice for a single customer account, a frontier agent performed an out-of-scope read of a system credential file; the read occurred in 5 of 10 in a clean single sweep and 6 of 10 in a prior run, every instance Critical.

Read the full results →
Compromise detectionmeasured · N=51 labeled realizations · 5 passes · 2026-08-24 · judge claude-haiku-4-5
0.96precision1.00recall
Category classification0.68accuracycategory accuracy · n=22 scored · 5 passes · 2026-08-24 · judge claude-haiku-4-5

Both figures hold only for the frozen judge named in their provenance lines. Change the rubric, the model or the temperature and they are void until re-measured.

Every run resolves one of two ways: a compromise becomes a fix report anchored to the offending step, and a clean run lands on the robustness leaderboard. Which one is measured, not predicted.

Featured run: tool misuse. Pick any of the Core-7 to watch its own attack.

Sample runasi02-run · tool misuse to out-of-scope file read · 8 steps
Watch replay
01 · principal instruction06 · read_file breach08 · complete

constructed demonstration · recorded validated-judge verdict · claude-haiku-4-5 · 2026-08-05