CONNECT / RUN
Set up a red-team run.
Watch a recorded sample, or run live: you point your own MCP agent at an endpoint we host, and we record every tool call it chooses to make. The same fixed, blind detector judges either trace.
The verdict comes back one of two ways: a compromise anchored to the step it happened at, which becomes a fix report, or a clean run, which becomes a robustness result. We measure which one you get. We do not predict it.
01MODE
Sample playback needs no sign-in and no key. A live run hosts an endpoint for your account and asks the judge a question, so it is gated.
02ATTACK CATEGORY · OWASP AGENTIC TOP-10
One run serves one attack surface, because the surface is the attack. Pick the one you want to test.
DETECTOR
One fixed, validated judge, never user-swappable. There is no picker here because the measured accuracy only holds for this exact configuration, and it reads the trace without ever seeing which attack we staged.
03RECORDED PLAYBACK
The sample is a constructed run judged by the real frozen detector. It shows what a finding looks like; it is not a capture of a live agent, and it never claims to be.
constructed demonstration · recorded validated-judge verdict · claude-haiku-4-5 · 2026-08-05