Fetching from the wire…
Models2026-09-21 · source-backed
Ben Swerdlow built a version of StarCraft playable only through agents and ran a full round-robin, publishing results September 19. Codex Astra at xhigh effort went 18-0, Astra at medium 16-2, Claude Fable 15-3, Astra at low 14-4, while Grok 4.6 at xhigh managed six command batches across 43 minutes. Styles diverged legibly, with Claude playing textbook macro and Codex favoring trick strategies. The ceiling is the finding, and it's the same shape as the CAD and ERP numbers: strong relative ordering among models, floor-level absolute performance against a human.
Each link below shares sources, entities, or timing with this story.
Posted to Show HN on September 4, it's a Rust loop engine that dispatches Claude, Codex, Hermes, Pi or NanoClaw against a codebase on a schedule, each run in a fresh isolated workbench inside a tmux session to prevent state leakage, with watchdog monitoring and REST, MCP and w...
Roland Gao published GoBench on September 15, scoring frontier models against a calibrated ladder of KataGo opponents. Astra Max 2,568, Astra High 2,227, Claude Opus 5 High 2,076, GPT-5.6 Sol Max 1,929, against KataGo's 4,400. Given coding tools and two hours of preparation be...
Two facts sit next to each other and neither cancels the other out. Anthropic published on September 4 that an internal general-purpose research model, roughly comparable to Claude Fable 5.1, formalized Fermat's Last Theorem in Lean over 11 days working largely autonomously. T...
Starting around 7:57 AM PT on September 3, all four reported outages simultaneously, with Downdetector logging 35,000+ US reports for ChatGPT, 1,400 for Claude and 1,200 for Grok before recovery by 12:38 PM PT. Cloudflare denied any significant disruption and xAI traced its ow...
The July 2026 update (v1.127-v1.131) runs each agent session against an isolated checkout, and it spans all three agents rather than being Copilot-only. Also: redesigned Agents window with side-by-side code review and chat, subagent tracking showing model, elapsed time and act...
UC Berkeley's Sky Lab put seven models through Claude Code, Codex CLI and Pi on 30 sampled tasks each from SWE-bench Lite and Terminal-Bench 2.0, three attempts per task, 21 model-harness pairs total. HarnessTax is the result, from Melissa Pan, Ion Stoica, Matei Zaharia and co...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.