Fetching from the wire…
Public story · 2026-07-25 · high
A nickel triages each session; real fixes escalate to a fifty-cent fixer, a split reused across 208 subagents.
Why now: The July 25 report is the first to attach session counts and dollar figures to the loop-engineering claims that circulated in June.
A developer's automated PR reviewer ran 1,785 agent sessions against real pull requests in five days, per a July 25 field report on dev.to.
The stakes are cost, not proof of concept. A five-cent triage agent screened each session. Only the runs that turned up something real moved to a fifty-cent specialist fixer, keeping most of the 1,785 sessions cheap.
The tool itself, a PR-babysitter, runs a five-state machine with 30-second polling. The same author ran a second job: five chained workflows and 208 subagents worked through roughly 600 archived sessions, 8GB of transcripts. The pass took two hours and fifty-three minutes and produced 1,883 findings, merged into 221 patterns. Eleven editor agents then used those patterns to rewrite the author's own skill library.
The report also restates an earlier data point: 64 Claude agents turned 535,496 lines of Zig into a 1,009,272-line Rust diff in eleven days. The bill ran to roughly $165,000. That project ran on a six-figure budget across 64 agents, a scale the PR-babysitter's nickel-and-dime triage split doesn't need.
Each link below shares sources, entities, or timing with this story.
The July 2026 update (v1.127-v1.131) runs each agent session against an isolated checkout, and it spans all three agents rather than being Copilot-only. Also: redesigned Agents window with side-by-side code review and chat, subagent tracking showing model, elapsed time and act...
July 22's board is plumbing, not apps: Kastra (199 votes) sells runtime authorization for Claude, Cursor, Codex, and OpenClaw with policy enforcement against prompt injection and unauthorized tool calls, while box (183 votes) sells Ubuntu VMs with SSH for agents at $0.036/hr (...
Rolling out July 7 through July 9, Cowork went mobile and web: sessions and files persist to the Claude account, background work continues after the laptop closes, scheduled tasks run with no device online, and approvals can be granted from your phone. Beta started with Max su...
Simon Willison surfaced Jarred Sumner's writeup of rewriting Bun's core from Zig to Rust this week, and the numbers stopped me cold. PR #30412, merged May 14, added roughly 1 million lines across 2,188 files, reached 99.8% test compatibility on Linux x64, and shrank the binary...
Anthropic commissioned the independent evaluator to test 72 injection scenarios, held out from Anthropic, each run 10 times against Fable 5, Opus 5, and Sonnet 5 as of July 17. Clean sweep. TechCrunch has the details. A third-party held-out eval is a much stronger claim than i...
1. Flip your multi-model pipeline to review-then-generate. Instead of using a reasoning model to plan before code generation, let the specialist generate freely and use reasoning tokens for review. Paper shows 90.2% pass@1 vs 87.2% for the planning pattern. Source 2. Audit you...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.