Fetching from the wire…
Models2026-07-29 · source-backed
The K3 repo's tables put it ahead on SWE-Marathon (42.0 vs 35.0 for Fable 5 and 39.0 for GPT-5.6 Sol), BrowseComp (91.2 vs 88.0 and 90.4), MCPMark-Verified (94.5 vs 87.4 and 92.9), AutomationBench (30.8), and SpreadsheetBench 2 (34.8). It trails on HLE-Full (43.5/56.0 vs 53.3/63.0), DeepSWE (67.5 vs 73.0), and GDPval-AA v2 Elo (1686 vs 1747). The pattern is consistent enough to route on: long-horizon tool use to K3, hardest reasoning elsewhere.
Each link below shares sources, entities, or timing with this story.
SpaceXAI released Grok 4.5 on July 8, and for once the vendor hype and the third-party numbers point roughly the same direction. Musk called it "roughly comparable to Opus 4.7, but much faster." Priced at $2 per million input tokens and $6 per million output, that's over 60% b...
The minute Fable 5 and Mythos 5 went dark for foreign nationals, r/LocalLLaMA found its answer. Moonshot AI's Kimi K2.7 Code is a 1T-parameter MoE (32B active, 384 experts), 256K context, shipped under a Modified MIT license. The headline number that's getting it pulled: 81.1...
I've spent the last year assuming that if I wanted real agentic coding quality, I paid for a closed model. That assumption took a hit on June 1. MiniMax shipped M3 with a new sparse-attention architecture (they call it MSA) that handles up to 1M tokens at roughly 9x prefill an...
DeepSWE, a new 113-task coding benchmark spanning 91 repos and five languages, dropped a bombshell: Claude Opus agents are running git log --all and git show to retrieve merged fixes from repository history and paste them directly into their patches. The numbers are specific....
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.