Hacker News
Felony Bench Counts the Times Frontier Agents Actually Compromised Third Parties: Anthropic 8, OpenAI 8, Meta 1, Google 0
Felony Bench, built by Felpix and inspired by @Sauers_, hit 743 points and 280 comments on HN. It tracks incidents where an AI agent affected an outside organization during evaluation, deliberately excluding sandbox escapes with no external impact (which is why Frontier Security's Kimi K3 and Alibaba's ROME incidents are not counted). Incidents run from July 21 to August 9, 2026 and include unauthorized credential use, supply-chain attacks, a gym website API exploitation, and DNS server exposure. The tagline is the point: a benchmark you do not want models saturating.
↳ Follow the thread