Fetching from the wire…
Public story · 2026-09-22 · source-backed
xAI released Grok 4.7 on September 21 on a new 2.1-trillion-parameter base, a 40% jump over Grok 4.6's 1.5T, with a 500K context window and pricing unchanged at $2 per million input and $6 per million output. Published scores: 46.3% on CursorBench 4.0 (up from 40.4%), 71.0% on DeepSWE v1.1 at high effort (up from 65.2%), 64.0% on EEBench, 19.6% on the Harvey legal agent benchmark. GitHub started rolling it into Copilot across every paid tier the same day. (xAI)
Then Artificial Analysis measured what it costs to get those scores. Grok 4.7 at xhigh effort scores 46 on the Intelligence Index using about 81,000 output tokens per task. Grok 4.6 at xhigh used 38,000. GPT-6 Astra at max uses 27,000. Time per task is about 7.1 minutes. Same rate card, double the bill. (Artificial Analysis)
This is the Harvey story again with the arrow pointing the other direction. Harvey's costs blew up because its own agent got chattier. Here the model got chattier and the price card didn't move, so a flat rate card hides a 2x cost increase from anyone who isn't measuring per-task.
The honest wins in the 4.7 release: hallucination rate dropped to 29% from 34%. Accuracy is flat at 47% against 48%. So you're paying twice as much per task for a model that's more careful about not making things up and no more correct overall.
Now put Xiaomi next to it. On September 21-22 Xiaomi released and open-sourced MiMo-V2.6-Pro, a 1.02T-total / 42B-active omnimodal MoE with a 1M context window, MIT-licensed weights on Hugging Face, at $0.435 per million input and $0.87 per million output. Artificial Analysis scored it at 46. The same composite number as Grok 4.7, at roughly a fifth the per-token price, from a phone manufacturer. (Latent Space)
Xiaomi also streamed six days of the RL run live on a public dashboard, including the failures: a Pro restart at step 17 from a GPU OOM caused by expert load imbalance, a grader-cluster network failure, and a cyber dataset removed after bad rollout patterns. Reported cost was $2,620,670 for Pro and $854,044 for Flash across 30 finished steps each. Those counters are self-reported and unauditable, and the 30 steps are the surviving tail of a longer job whose discarded compute appears nowhere. Still, nobody else is publishing their OOM restarts. (Traictory)
The HN launch thread for Grok 4.7 ran 582 points and reads like a split decision. Commenter moojacob points out the model carries 40% more weights at identical pricing and calls the delayed release disappointing ahead of Opus 5.5. Simon Willison ran his SVG test and got a bicycle with the seat and pedals in the wrong places. Praise clusters on parallel tool calls and frontend work. (Hacker News)
Stop reading rate cards as cost. Run your own ten hardest tasks through any model you're evaluating and record total output tokens and wall clock, then divide.
Each link below shares sources, entities, or timing with this story.
The Astra coverage went to price and context window. The number that changes how I'd deploy it went into a system card nobody read. Artificial Analysis measured GPT-6 Astra's hallucination rate on AA-Omniscience at 51% at max effort, against 92% for its predecessor. Accuracy w...
GitHub published Project HydraFusion on September 4. Spotify published Portal on September 3. CodeRabbit published its Astra evaluation on September 4. None of them coordinated, and all three are the same argument. HydraFusion is a Copilot research preview that treats workflow...
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
Artificial Analysis ranked it 46 on the Intelligence Index, #1 of 114, ahead of Kimi K3 at 44 and GLM-5.3 at 45, with leading closed models at 53. Natively omnimodal MoE, 1M context, weights on Hugging Face, $0.435 per million input and $0.87 per million output. The companion...
Anthropic released Claude Fable 5.1 on September 1. Claude Code v2.1.257 made it the default Fable model at 17:53 UTC that day, with a 1M-token context window, $10 per million input tokens, $50 per million output, and $0.25 per million on cache reads (claude-code CHANGELOG). B...
xAI shipped it August 12 with a 500K context, February 2026 cutoff, $2/$6 per million. It scored 61 on the Artificial Analysis Index, tying GPT-5.6 Sol Max, one point behind Fable 5 Max. The number that got 334 points and 381 comments on HN is from Artificial Analysis's teardo...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.