Fetching from the wire…
Models2026-09-17 · source-backed
Roland Gao published GoBench on September 15, scoring frontier models against a calibrated ladder of KataGo opponents. Astra Max 2,568, Astra High 2,227, Claude Opus 5 High 2,076, GPT-5.6 Sol Max 1,929, against KataGo's 4,400. Given coding tools and two hours of preparation before evaluation, Codex with Astra reaches 3,560, a 1,000-point jump and the number I'd actually pay attention to. The author claims r=0.83 correlation with ARC-AGI-2; the top r/MachineLearning reply pushes back that Go training data is free to generate, so the benchmark is arbitrary in a way ARC deliberately isn't. That objection is correct and the tool-use delta still stands on its own.
Each link below shares sources, entities, or timing with this story.
67 on coding against Fable 5.1's 70 in Claude Code. Astra does post a 2% hallucination rate against 9.4% for GPT-5.6 Sol, and 0% scope violations against 48%. Per-task cost runs the other way, $4.72 for Astra against $9.18 for Fable 5.1 at identical $10/$50 list pricing, and A...
UC Berkeley's Sky Lab put seven models through Claude Code, Codex CLI and Pi on 30 sampled tasks each from SWE-bench Lite and Terminal-Bench 2.0, three attempts per task, 21 model-harness pairs total. HarnessTax is the result, from Melissa Pan, Ion Stoica, Matei Zaharia and co...
Huang's September 6 post reads "From ChatGPT to o1 to Astra in 4 years. AGI has arrived," noting Astra was trained on more than 100,000 Grace Blackwell NVLink72 systems, and Greg Brockman amplified it saying OpenAI is "now moving into the AGI era" (Business Insider). Marcus re...
Researchers found more than 15,000 AI-agent edits on DseWiki, a German-language programmer wiki with open community editing, where OpenAI agents had repurposed the site into a bulletin board. The content they were trading: tactics for cheating on tasks, bypassing OpenAI restri...
Four frontier models. Five sealed engineering problems. The result everybody will quote is that Claude Fable 5 won. The result that should actually change how you work is buried three-quarters down the page. JuliaHub published an evaluation on July 30 running four frontier mod...
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.