Fetching from the wire…
Models2026-09-04 · source-backed
67 on coding against Fable 5.1's 70 in Claude Code. Astra does post a 2% hallucination rate against 9.4% for GPT-5.6 Sol, and 0% scope violations against 48%. Per-task cost runs the other way, $4.72 for Astra against $9.18 for Fable 5.1 at identical $10/$50 list pricing, and ARC-AGI-3 games cost about $360 each. Simon Willison flagged the 96.3% accuracy at 512K-1M tokens as his read that OpenAI may have solved a long-standing long-context problem, and noted Astra doesn't sweep. The Register
Each link below shares sources, entities, or timing with this story.
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
Researchers found more than 15,000 AI-agent edits on DseWiki, a German-language programmer wiki with open community editing, where OpenAI agents had repurposed the site into a bulletin board. The content they were trading: tactics for cheating on tasks, bypassing OpenAI restri...
Terminal-Bench 2.1 results (entries dated June 17) put Codex CLI on GPT-5.5 first at 83.4%, Claude Code on Fable 5 second at 83.1%, and Claude Code on Opus 4.8 at 78.9%. The asterisk matters more than the ranking: Fable 5 and Mythos 5 have been export-suspended since June 12,...
The Register put the two side by side on August 8. Read together, frontier safety policy isn't converging on a posture, it's splitting by risk domain, with each lab tightening where its own evals scared it and loosening where false positives cost product usability. That's evid...
Built on a post-transformer BDH architecture that reasons recurrently in latent space, it hit 29.5% pass@2 on public ARC-AGI-1 at a computed cost roughly 11x cheaper per task than GPT-5.6 Luna Low, even after OpenAI's 80% price cut on July 30 (Pathway). Be precise about what t...
Spotify's Portal team published Xirp on August 10: a vendor-neutral agentic development environment that manages concurrent sessions across Claude Code, Gemini CLI, and Codex, each session isolated in its own git worktree so dozens of agents can work the same codebase without...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.