Fetching from the wire…
Public story · 2026-07-19 · high
The open-weight model beats Anthropic's Opus 4.8 by one point but trails Fable 5 by three, and ties two GPT-5 variants on coding.
Why now: AINews published the placement numbers in its July 19 briefing.
Kimi K3 scored 57 on the Intelligence Index, one point ahead of Anthropic's Opus 4.8, per AINews' numbers in Latent Space's July 19 briefing.
That edge matters for anyone deciding whether to build agents on closed or open-weight models. A downloadable model outscoring a flagship closed one, even narrowly, weakens the argument that closed weights are worth their premium.
K3 still trails Fable 5, which leads the Intelligence Index at 60, by three points. Its Coding Agent Index score of 57 ties GPT-5.6 Terra and GPT-5.5 rather than beating them. On DeepSWE, K3 ranks third among open-weight models, not first.
AINews' numbers don't say what K3 costs to run at Opus-class throughput, or how it performs on anything outside these three benchmarks. A one-point lead on an aggregate score doesn't cover long-context reliability or tool use in a live agent loop.
A downloadable model narrowly outscoring Opus on this index means the closed-open frontier gap is now single points, not tiers. Whether K3 holds its DeepSWE ranking against the next open-weight release is the one to watch.
Each link below shares sources, entities, or timing with this story.
SpaceXAI released Grok 4.5 on July 8, and for once the vendor hype and the third-party numbers point roughly the same direction. Musk called it "roughly comparable to Opus 4.7, but much faster." Priced at $2 per million input tokens and $6 per million output, that's over 60% b...
Everyone kept score wrong. When OpenAI shipped GPT-5.6 (the Sol flagship plus Terra and Luna) to GA on July 9, then xAI put out Grok 4.5, Meta dropped Muse Spark 1.1, and Cognition shipped SWE-1.7, the reflex was to ask who won the benchmark. Wrong question. On the Artificial...
Everyone benchmarks per task. Accuracy on SWE-bench, pass rate on Terminal-Bench, a leaderboard row per model. Together AI ran the experiment sideways: fix the budget at $100, point both models at DeepSWE, and count how much work came out the other end. GLM-5.3 finished 17 tas...
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
AINews resumed publishing after its post-Kimi-K3 blackout with a "not much happened today" edition (Latent Space). That's a real signal after GPT-5.6 Sol, Grok 4.5, Meta Muse, Kimi K3, and the Qwen3.8-Max preview all landed inside a fortnight. Matthew Berman published another...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.