Boundary-Bench: turn on the security controls your enterprise already runs and coding agents lose up to 18.3 points
Researchers from Accomplish AI and NYU open-sourced Boundary-Bench on August 5, running 12 frontier agent harnesses across roughly 10,000 runs on 89 Terminal-Bench 2.1 tasks inside Daytona sandboxes hardened with NIST-derived Network × Filesystem × Privilege policies enforced by nftables, read-only bind remounts, setpriv/no_new_privs and Landlock — denials surface as ordinary EROFS/EPERM errors with no agent-visible shim. Codex on GPT-5.6 Sol leads unrestricted at 83.9%, but under the strictest tier Grok Build on Grok 4.5 tops out at 74.9% and Claude Code on Sonnet 5 shows the largest degradation at 18.3 points; costs inflate up to 167.3%. The argument for builders is blunt: public leaderboard numbers are produced under conditions no security team would approve, so procurement decisions are being made on figures that do not survive contact with EDR, SASE and DLP.
↳ Follow the thread