Fetching from the wire…
Vibe Coding2026-07-29 · source-backed
On the May 15 – July 1 window (111 problems from 65 repositories, all post-dating training cutoffs), per the announcement thread and the leaderboard itself: Fable 5 at 64.5%, Grok 4.5 at 63.8%, Opus 5 at 63.4%, GLM-5.2 at 62.9%, GPT-5.6 Sol at 62.3%. Compare that to the 95.0 Fable 5 posts on SWE-bench Verified. On genuinely unseen repos the frontier is basically a coin flip, so route on cost and latency instead of leaderboard rank. A multilingual slice covering Go, Java, Rust, and TypeScript is new but single-sourced to the announcement.
Each link below shares sources, entities, or timing with this story.
SpaceXAI released Grok 4.5 on July 8, and for once the vendor hype and the third-party numbers point roughly the same direction. Musk called it "roughly comparable to Opus 4.7, but much faster." Priced at $2 per million input tokens and $6 per million output, that's over 60% b...
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
Everyone kept score wrong. When OpenAI shipped GPT-5.6 (the Sol flagship plus Terra and Luna) to GA on July 9, then xAI put out Grok 4.5, Meta dropped Muse Spark 1.1, and Cognition shipped SWE-1.7, the reflex was to ask who won the benchmark. Wrong question. On the Artificial...
This is the other half of the Fable 5 story, so read them together. While the best coding model in the world is uncallable, an open-weight one quietly posted frontier-adjacent numbers. Per Tom's Hardware, independent benchmarks for the MIT-licensed GLM-5.2 (744B params, 40B ac...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Zhipu's GLM-5.1 ranked third on Code Arena, jumping 90+ points over its predecessor GLM-5 and landing ahead of GPT-5.4 and Gemini 3.1 Pro. Two separate r/LocalLLaMA threads (491 upvotes and 233 upvotes) confirm this isn't just a benchmark curiosity. Practitioners are paying at...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.