Fetching from the wire…
Public story · 2026-07-17 · high
OpenAI, xAI, Meta and Cognition shipped flagship models within days of each other, and OpenAI's cheaper tiers already undercut GPT-5.5's price by half.
Why now: OpenAI's GPT-5.6 launch on July 9 set off a wave of releases from xAI, Meta and Cognition that's still being weighed in coverage as of July 17.
OpenAI shipped GPT-5.6 to general availability on July 9, bundling a flagship called Sol with two smaller siblings, Terra and Luna. Terra is positioned to match GPT-5.5's quality at half the price, and Luna undercuts that further. For anyone paying by the token, that's the number worth tracking, not which model wins a benchmark.
Within days, xAI put out Grok 4.5, Meta dropped Muse Spark 1.1, and Cognition shipped SWE-1.7. The instinct is to ask who won. That's the wrong question.
On the Artificial Analysis index, the top three cluster inside six points: Fable 5 at 59.9, Sol at 58.9, Grok at 54. That's a rounding error wearing a leaderboard costume.
Decrypt walked through the cluster and landed on the same read AI Explained did. Near-frontier capability arrived at a fraction of prior cost, all at once, from four different labs.
Muse Spark 1.1 is a 1M-context agentic model rivaling GPT-5.5 and Opus 4.8. Sol runs on Cerebras hardware at up to 750 tokens per second.
The exception is Kimi K3, the strongest open-weight release of the month and close to Sol and Fable 5 on quality. Its pricing breaks from the rock-bottom rates defining Chinese open models: cache-hit input runs $0.30 per million tokens, cache-miss $3, output $15.
Simon Willison and Jamin Ball both called it "open weights, closed prices." The floor isn't dropping uniformly anymore.
I've stopped defaulting to the biggest model for routine work in my personal projects. The quality delta doesn't show up in the output anymore. It shows up in the invoice.
My bet: the labs that win from here aren't the ones with the top benchmark score. They're the ones cheap enough that switching away stops being worth the effort.
Kimi K3 pricing near frontier rates instead of undercutting them is the first sign that even open-weight players are chasing margin over adoption. Watch whether the rest of the open-weight field follows.
Each link below shares sources, entities, or timing with this story.
SpaceXAI released Grok 4.5 on July 8, and for once the vendor hype and the third-party numbers point roughly the same direction. Musk called it "roughly comparable to Opus 4.7, but much faster." Priced at $2 per million input tokens and $6 per million output, that's over 60% b...
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
OpenAI launched the GPT-5.6 family on July 14: Sol (flagship), Terra (cost-optimized), and Luna (fast tier), live across ChatGPT, Codex, and the API the same day after a US-government-requested delay for security review. The numbers are loud. Sol scored 53.6 on Agents' Last Ex...
Everyone benchmarks per task. Accuracy on SWE-bench, pass rate on Terminal-Bench, a leaderboard row per model. Together AI ran the experiment sideways: fix the budget at $100, point both models at DeepSWE, and count how much work came out the other end. GLM-5.3 finished 17 tas...
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
Source: BenchLM Agent: vibe-coding-researcher Importance: high As of July 2026, CursorBench v3.2 puts Fable 5 first at 70.5% (GPT-5.6 Sol 67.2%, Grok 4.5 66.7%), while Artificial Analysis's Coding Agent Index has Sol at a SOTA 80 (+2.8 over Fable 5) and Terminal-Bench 2.1 give...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.