Your agent works in the demo. We run it through adversarial scenarios mapped to the OWASP Top 10 for LLM Applications — so the failure modes show up in a sandbox, not in an incident.
Six categories from the OWASP Top 10 for LLM Applications (2025), the closest thing this space has to a standard risk checklist.
Not covered yet: LLM03 supply chain and LLM04 data/model poisoning — those are governance problems, not runtime testing problems. Honest scope, not a longer feature list.
This is a scripted preview of the sandbox output — pick any card below and run it.