Fetching from the wire…
Public story · 2026-07-19 · high
Berman's July 18 video found the win limited to Arena's Frontend Code eval, with K3 placing third on GDPval-AA v2.
Why now: Berman's video landed two days after Moonshot's July 16 release, while the benchmark claims were still spreading.
Matthew Berman spent 12 minutes testing Kimi K3's benchmark claims against Moonshot's own numbers, two days after the July 16 release.
The headline win holds for one eval, Arena's Frontend Code test, while K3 placed third on the broader GDPval-AA v2. That's a category score, not a general one, and it's the number anyone choosing between K3 and Fable on coding work should weigh.
Frontend code is one skill among many an eval suite measures. GDPval-AA v2 is the broader test, and third place there didn't make Moonshot's July 16 headline.
The video works through Moonshot's category-by-category numbers instead of the launch post's framing. It pulled around 72,000 views.
Kimi K3's Frontend Code win will keep getting cited as a general benchmark victory. Most coverage repeats a headline number instead of checking what an eval suite actually measures. Watch whether Moonshot's next release leads with the same single-category framing.
Each link below shares sources, entities, or timing with this story.
AINews resumed publishing after its post-Kimi-K3 blackout with a "not much happened today" edition (Latent Space). That's a real signal after GPT-5.6 Sol, Grok 4.5, Meta Muse, Kimi K3, and the Qwen3.8-Max preview all landed inside a fortnight. Matthew Berman published another...
Moonshot AI released Kimi K3, a sparse mixture-of-experts activating 16 of 896 experts per token. That's about 1.8% of the pool live at any moment, with a 1M-token context window and native vision. Two new architectural pieces show up: Kimi Delta Attention and Attention Residu...
paddo.dev makes the most contrarian read: the letter's substance isn't openness but paragraph nine, defending distillation as "a widely used technique for model improvement" and urging policymakers against "conflating legitimate model development techniques with misappropriati...
Everyone kept score wrong. When OpenAI shipped GPT-5.6 (the Sol flagship plus Terra and Luna) to GA on July 9, then xAI put out Grok 4.5, Meta dropped Muse Spark 1.1, and Cognition shipped SWE-1.7, the reflex was to ask who won the benchmark. Wrong question. On the Artificial...
Anthropic commissioned the independent evaluator to test 72 injection scenarios, held out from Anthropic, each run 10 times against Fable 5, Opus 5, and Sonnet 5 as of July 17. Clean sweep. TechCrunch has the details. A third-party held-out eval is a much stronger claim than i...
The model docs page hit 457 points on Hacker News for a K3 variant capped at 256k context that Moonshot says delivers the same results within that window at roughly half the quota of the 1M K3. Positioned for everyday Q&A, code completion, routine feature work, and single-file...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.