Fetching from the wire…
Public story · 2026-08-09 · high
Boundary-Bench ran 12 agent harnesses through real firewall and filesystem locks, and costs climbed as much as 167 percent as those restrictions tightened.
Why now: Boundary-Bench's code and full run data landed on GitHub on August 5, giving buyers their first hardened-environment comparison across 12 agent harnesses.
Claude Code on Sonnet 5 loses 18.3 points, the steepest fall among the 12 agents Boundary-Bench tested under enterprise security policy, per its GitHub release.
That gap matters for procurement teams. They're picking coding agents off leaderboard numbers that assume no firewall, no filesystem lock, no privilege controls. Turn those on, and the field compresses. Codex on GPT-5.6 Sol leads the unrestricted board at 83.9%. Under the strictest security tier, the best score drops to 74.9%, from Grok Build on Grok 4.5.
Boundary-Bench comes from Accomplish AI and NYU, open-sourced on GitHub on August 5. The project ran roughly 10,000 test runs across those 12 harnesses on 89 Terminal-Bench 2.1 tasks, inside Daytona sandboxes.
It combines NIST-derived network, filesystem and privilege policies enforced with nftables, read-only bind remounts, setpriv/no_new_privs and Landlock. Those are controls modeled on what a real security team runs. When an agent hits a blocked path, it just gets a plain EROFS or EPERM error. No shim tells it a policy stopped it.
Cost climbs too. Boundary-Bench measured token costs inflating by as much as 167.3% as security restrictions tighten. Its release doesn't explain why, only that the number climbs with policy strictness.
Related coverage: NVIDIA's SkillSpector scanner found 26.1% of 42,447 agent skills carry a vulnerability, a sign that agent security gaps extend beyond sandbox policy.
Public leaderboard numbers stop being valid for procurement the moment a team's own firewall and filesystem policy gets applied to the agent. Watch whether vendors start publishing their own hardened-tier scores, or whether Boundary-Bench's ranking becomes the one buyers actually trust.
Each link below shares sources, entities, or timing with this story.
Same source domain / Semantically similar
Reported by the same outlet (github.com); covers closely related ground (similarity 0.76).
Reported by the same outlet (github.com); covers closely related ground (similarity 0.76).
Same source
Cite the same source (GitHub (boundary-bench)).
Cite the same source (GitHub (boundary-bench)).
Same source domain / Semantically similar
Reported by the same outlet (github.com); covers closely related ground (similarity 0.75).
Reported by the same outlet (github.com); covers closely related ground (similarity 0.74).
Reported by the same outlet (github.com); covers closely related ground (similarity 0.73).
Reported by the same outlet (github.com); covers closely related ground (similarity 0.72).