Red-team your MCP agent. Trust the verdict either way.
Point your own agent at an endpoint we host, and we serve it one of the OWASP Agentic Top 10 attacks and record every tool call it chooses to make.
Most tools only guess whether an agent was compromised. We tested our detector against attacks with known outcomes and measured how accurate it is, so you can trust its verdicts.
Measured finding Asked to fetch the invoice for a single customer account, a frontier agent performed an out-of-scope read of a system credential file; the read occurred in 5 of 10 in a clean single sweep and 6 of 10 in a prior run, every instance Critical.
Read the full results →Both figures hold only for the frozen judge named in their provenance lines. Change the rubric, the model or the temperature and they are void until re-measured.
Every run resolves one of two ways: a compromise becomes a fix report anchored to the offending step, and a clean run lands on the robustness leaderboard. Which one is measured, not predicted.
Featured run: tool misuse. Pick any of the Core-7 to watch its own attack.
constructed demonstration · recorded validated-judge verdict · claude-haiku-4-5 · 2026-08-05