Fetching from the wire…
Models2026-09-21 · source-backed
The Center for AI Safety's index added Fable and Astra entries, with the live dashboard showing 20.83% full automation. RLI grades models on real freelance work priced by the humans who did it: game development, product design, architecture, data analysis, video animation, over 6,000 hours and $140,000 of projects, some costing over $10,000 and taking 100+ hours, judged end-to-end from a single prompt by human experts. Astra fully completed about one project in five. Reaching 90% counts as a fail. The harshness is the design.
Each link below shares sources, entities, or timing with this story.
Ara Kharazian, who runs the Ramp AI Index, posted September 16 that OpenAI's growth is coming from shifts off GPT-5.6 Sol, off some Anthropic models, and from net-new usage, reading it as Anthropic's pace-the-frontier position costing it frontier adoption. That cuts against th...
Vittorio Romeo posted working examples: a Prey (2017) gunplay overhaul adding aim-down-sights and weapon-versus-environment collision to a game that never supported ironsights, and a patch decoupling Touhou 11 and 12 from their hardcoded 60 FPS game-logic assumption so they ru...
67 on coding against Fable 5.1's 70 in Claude Code. Astra does post a 2% hallucination rate against 9.4% for GPT-5.6 Sol, and 0% scope violations against 48%. Per-task cost runs the other way, $4.72 for Astra against $9.18 for Fable 5.1 at identical $10/$50 list pricing, and A...
Every number you use to pick a harness comes from public repositories the models may have trained on. Specific Labs built the version that doesn't: tasks drawn from licensed private company codebases, including a 200K-user event app and a fintech processing over 100K bank stat...
His September 12 post calls Astra a bigger jump than Fable 5 to 5.1, citing 98% on FrontierMath Tier 4, 98.1% on extended NYT Connections against Fable 5.1's 90%, the first autonomous Montezuma's Revenge clear and a one-shot Portal completion. The monitorability problem is Nee...
Ollama cut v0.34.0-rc1 on September 5 at 23:49 UTC, and the headline item changes the shape of the local-versus-hosted decision rather than the performance of either side: Ollama-hosted open models can be selected directly inside ChatGPT Desktop, with setup driven from the Oll...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.