Fetching from the wire…
Tools2026-09-07 · source-backed
Registry data confirms 2.1.259 (Sep 2), 2.1.260 (Sep 3), 2.1.261 (Sep 4) and 2.1.263 (Sep 6), with no 2.1.262 ever published (npm dist-tags, via Creative AI News). Together they add unattended-run plumbing: managedMcpServers, --permission-prompts none, /reload-plugins in headless sessions, /skill-doctor, and bashOutputMaxChars/taskOutputMaxChars raised to 128K. The catch is in the dist-tags: stable reads 2.1.236 against latest at 2.1.263, so anyone pinned to stable is missing all of it including the 2.1.259 and 2.1.260 permission-enforcement fixes for Read() deny-rule gaps and a zsh command-substitution bypass.
Each link below shares sources, entities, or timing with this story.
67 on coding against Fable 5.1's 70 in Claude Code. Astra does post a 2% hallucination rate against 9.4% for GPT-5.6 Sol, and 0% scope violations against 48%. Per-task cost runs the other way, $4.72 for Astra against $9.18 for Fable 5.1 at identical $10/$50 list pricing, and A...
Context Privilege Escalation names two classes, M-CPE where attacker-controlled low-privilege content gets folded into a higher-privileged message role, and X-CPE where it persists past the context that introduced it. The authors ran it against 12 production harnesses includin...
Rust, created May 14, at 2,643 stars (GitHub). Every run produces checkpoints linking a commit to the session that made it, including prompts, tool calls and reasoning. It runs Claude Code, Codex, its own agent and anything from the ACP registry side by side against one codeba...
A free-alpha macOS app plus GitHub extension and MCP server that answers "why" questions from a repo's own pull requests and issues, shows the evidence, and says nobody wrote this down when nobody did (Icarus). The insight underneath: merged PRs leave commits but refused ones...
Raj Nagulapalle's FetchSandbox MCP took 107 votes on August 23, wiring 70+ API sandboxes into Cursor or Claude Code via MCP config. The claim is narrower and more testable than most agent tooling: reproduce the real integration failure against a sandbox, apply the fix, re-run...
Deng et al. built 120 real-case-grounded tasks across 20 business scenes in six financial domains, running four self-evolving scaffolds on a shared Qwen3.7-Max backbone against paired non-evolving controls. Letta posted the highest evolved score (91.65) and fewest compliance i...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.