Fetching from the wire…
Public story · 2026-09-09 · high
V4 Pro launched August 13 and is already being redirected to a cheaper model that beats it on cost and matches it on math benchmarks.
Why now: DeepSeek's routing change takes effect around September 10 Beijing time, the same window three other releases pushed inference cost down again.
DeepSeek told users that once V4.1 Flash launches around September 10 Beijing time, every V4 Pro API request will route to V4.1 Flash instead and bill at Flash pricing, according to a r/LocalLLaMA thread. V4 Pro reached general availability August 13. Three weeks as the flagship, then a redirect to a cheaper model with no deprecation notice and no migration guide.
The numbers back up the swap. V4.1 Flash scores 82.7 on Terminal-Bench 2.1 at $0.14/$0.28 per million tokens, and on MathArena's AIME 2026 set it hits 95.83% against Pro's 96.67%, per The New Stack. Statistically indistinguishable math performance at about a ninth of the per-problem cost, running around 400 tokens per second. New pricing from September 10 drops the input cache-hit cost to $0.003 per million tokens.
DeepSeek had already opened an intermediate build to all API users on September 8 under the model ID deepseek-v4.1-flash-expires-on-0910, a self-destructing identifier good for exactly two days, no beta signup required, per another r/LocalLLaMA post.
Three more releases went up in the same 48 hours, chasing the same cost line down. Inception Labs released Mercury 2.5, a diffusion model running 1,107 tokens per second, discounted 80% to $0.04/$0.15 per million at Inception Labs. Desert Ant Labs gave away 18 on-device models free to 100,000 monthly devices with no per-token metering, according to its launch post. And a GitHub project called deltafin got the 2.8-trillion-parameter Kimi K3 running on a MacBook Pro by streaming 1.45TB of weights off four SSDs.
Pin model IDs explicitly in every client. DeepSeek just proved a vendor will redirect a paid tier to a different model while your evals still point at the old one. And if your pricing assumed a fixed cost per AI call, that assumption is now wrong every few months, not every few years.
Each link below shares sources, entities, or timing with this story.
DeepSeek-V4-Flash-0731 landed July 31 under MIT with a DSpark speculative-decoding module attached. Terminal Bench 2.1: 82.7. Toolathlon-Verified: 70.3. DSBench-FullStack: 68.7. DeepSWE: 54.4. NL2Repo: 54.2. The model card claims it beats DeepSeek-V4-Pro (Preview) "despite its...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Writer launched Palmyra X6 on August 13 with a number that should reset how you think about agent COGS: 52% lower average cost, 48% better speed, 10% better quality. The model is a post-training variation of Z.ai's open-source GLM-5.2. A US enterprise SaaS vendor built its fla...
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
DeepSeek released V4 on April 24 and the numbers demand attention. V4-Pro is 1.6 trillion parameters total with 49 billion active, MIT-licensed, native 1M-token context. It scores 80.6% on SWE-bench Verified, putting it within 0.2 points of Claude Opus 4.6. On Terminal-Bench 2...
Per-token prices went down at both labs. Subscriptions are draining faster at both labs. Those aren't in tension once you look at token counts. On the OpenAI side, r/OpenAI collected reports from Linux.do and NodeSeek alleging Astra consumes more Plus quota than its published...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.