Fetching from the wire…
Public story · 2026-09-12 · high
Accomplish found sandbox escapes in Claude Code, Codex and Cursor; Anthropic's fix took the longest of the three.
Why now: Reported in coverage dated September 12, 2026.
Sandbox escapes in Claude Code, Codex and Cursor let coding agents break out of their containers and reach systems outside them, according to Accomplish founders Amit Avner, Or Hiltch and Guy Zipori, as reported by Upstarts Media.
The sandbox is what stands between an agent's mistake, or a prompt injection, and the actual machine it's running on. A leak there isn't cosmetic.
Cursor fixed its issue in about a week after a July report. OpenAI fixed two separate issues in August on a similar timeline and confirmed both on the record. Anthropic's fix took roughly 50 days and about 30 releases to land, and Anthropic declined to comment. Cursor also declined to comment on the record.
That's roughly 7x longer than the other two vendors, going by the report's own numbers. Neither company has said why. Thirty releases is a lot of surface for one patch to move through, and the piece doesn't say whether the gap was triage, complexity, or something else entirely.
Accomplish has a commercial interest in this finding, per the report. That doesn't make the bugs fake, but it's a reason to want the timeline confirmed independently rather than taken at face value.
Ask a vendor how fast they patch a sandbox escape once one's confirmed, not just whether one exists. On this report, the answers aren't the same.
Each link below shares sources, entities, or timing with this story.
Six clients. One manifest. Zero vendor lock. Vercel published Agent Plugins 1.0.0 on August 6, an openly licensed spec that bundles Agent Skills and MCP servers behind a single portable manifest. The shape is deliberately boring: a plugin.json requiring only schemaVersion and...
The IDE market is fragmenting, and this week drew the sharpest lines yet. Cursor 3 launched as a rebuilt agent-orchestration platform in Rust and TypeScript, replacing the VS Code fork with an Agents Window for dispatching and monitoring multiple AI coding agents. Anysphere hi...
Three things happened this month that only make sense together. Agent Plugins 1.0 shipped co-signed by six competitors: AWS, Anysphere, Microsoft, OpenAI, Vercel and Google (GitHub Changelog). It makes skills-plus-MCP bundles portable across clients. OpenAI's August 11 Codex c...
Three frontier models shipped in a single week this month, and teams with a standing eval harness had a routing decision in hours. Anthropic's own agent-eval guidance says 20-50 tasks drawn from your real usage and real failures is enough to detect issues (DeepEval). DeepEval...
Simon Willison spent a while taking ChatGPT Work apart and published the map on August 30. Work splits into Work Cloud and Work Local, the latter being the renamed Codex desktop app, at $20/month and up since July 9. He enumerates six capabilities Work has that Chat doesn't, a...
CherryHQ/cherry-studio (50,068 stars) cut v2.0.0 on August 5 and reached 2.0.2 by August 7, rebuilding from chat client into a work surface with multi-window and split-view layouts. The release notes say the subscription part explicitly: OpenAI and Anthropic endpoints, reusing...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.