StealthBench Finds No Model Exceeds a 54% Safe Success Rate: Offensive Agents Finish the Job but Blow Their Own Cover
Ads Dawson and Adrian Wood's July 28 benchmark distills 11 verified real security incidents into 14 dockerized scenarios and scores autonomous offensive agents across six OPSEC dimensions using a three-model LLM judge panel with majority vote. Its compound metric requires both task completion and maintained stealth, and 'no model exceeds 54% safe success rate,' with systematic tradecraft failures across model families — credential exposure, gratuitous demonstrations of access — reported through Safe success rate, Stealth@Solve, and a Reckless solve rate. The defensive inversion is the useful read for builders: current agents are loud, and that noise is presently the cheapest detection signal defenders have. Benchmark, leaderboard, and dataset are public.
↳ Follow the thread