Fetching from the wire…
Public story · 2026-08-07 · high
The update reuses 4.5's 1.5T-parameter base and credits training tweaks, not scale, while a 2.1T Grok 4.7 looms weeks out.
Why now: This lands in the August 7 briefing cycle, days before a much bigger Grok 4.7 is expected to arrive.
xAI shipped Grok 4.6 on August 7 with no model card, no benchmark scores, and no way to verify what changed, per reporting from Evolink.ai. That leaves developers guessing: every performance number attached to 4.6 is speculation, not measurement, until independent Arena testing produces a score.
The model reuses the same 1.5 trillion-parameter V9 foundation that powered Grok 4.5. xAI is crediting the gains to improved supervised fine-tuning and reinforcement learning, not a bigger base model, according to the same report. That's a meaningful distinction if true: whatever 4.6 does better, it's doing with the same raw compute as its predecessor.
For context, Grok 4.5 in its "high" configuration scores 54 on the Artificial Analysis Intelligence Index, putting it fourth behind Claude Fable 5, GPT-5.5, and Claude Opus 4.8. Whether 4.6 moves that number up, down, or stays flat is unknown until someone runs the tests xAI didn't publish.
A 2.1T-parameter Grok 4.7 is reportedly weeks away, per the same August 7 report.
Shipping a mid-cycle model with zero verification while a much bigger successor is already queued reads like a placeholder move, not a flagship claim. If 4.6 mattered as a capability jump, xAI had every reason to publish the numbers on release day. Watch whether Arena scores show up before 4.7 lands. If they don't, 4.6 was never meant to be judged on its own.
Each link below shares sources, entities, or timing with this story.
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Moonshot AI released Kimi K3, a sparse mixture-of-experts activating 16 of 896 experts per token. That's about 1.8% of the pool live at any moment, with a 1M-token context window and native vision. Two new architectural pieces show up: Kimi Delta Attention and Attention Residu...
Launched August 8 as the new Quality Mode at grok.com/imagine and in the Grok mobile apps, pitching precision editing, crisp text rendering, and improved factuality, with API access promised but not shipped (The Decoder). On the August 7 Arena leaderboards the faster "low" var...
SpaceXAI released Grok 4.5 on July 8, and for once the vendor hype and the third-party numbers point roughly the same direction. Musk called it "roughly comparable to Opus 4.7, but much faster." Priced at $2 per million input tokens and $6 per million output, that's over 60% b...
Four frontier models. Five sealed engineering problems. The result everybody will quote is that Claude Fable 5 won. The result that should actually change how you work is buried three-quarters down the page. JuliaHub published an evaluation on July 30 running four frontier mod...
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.