Fetching from the wire…
Public story · 2026-07-20 · high
Moonshot split Kimi into separate chat and coding subscriptions, exposing what agentic work costs to serve.
Why now: Moonshot suspended signups on July 20, the same day an unsealed 2022 Altman email surfaced reframing open-weight releases as a competitive weapon rather than a gift.
Moonshot AI suspended new consumer subscriptions to Kimi K3 on July 20, about 48 hours after launch, because request volume maxed out its compute cluster. Remaining GPUs now go to existing paid subscribers only. TechNode broke the story, and Reuters, SCMP, and Global Times corroborated it, so this isn't a rumor.
The fix splits membership into two products: a general Kimi Membership and a separate Kimi Code Membership for programming. That split is an admission.
Coding tasks make dozens or hundreds of model calls per task, each carrying a large context window, versus one call for a chat turn. Moonshot can't serve that at general-purpose rates.
The same July 20 coverage lines up with two more data points. Anthropic's Bun migration burned 5.9 billion input tokens against 690 million output tokens. OpenAI set Codex's context cap right at its 2x price break.
Three companies, three different responses to the same problem: pricing hasn't caught up to what agentic coding costs to serve.
Two days before the Kimi suspension, Alibaba previewed Qwen3.8-Max at WAIC Shanghai. It's a claimed 2.4 trillion parameters, multimodal, priced at 10% of standard rates through Token Plan and Qoder.
There's no benchmark table, no model card, and no disclosed active-parameter count. For a sparse mixture-of-experts model, that's the one number that says what it actually costs to run, and it's not an oversight.
Simon Willison surfaced an unsealed 2022 email from Sam Altman to OpenAI's board, part of the Musk v. Altman discovery. In it, Altman proposed shipping a GPT-3-class open model to make funding harder for rival labs, not to share the technology.
Open-sourcing as a competitive weapon, on the record. The same week, OpenAI's policy lead called Chinese open models "full AI communism."
Kimi K3's weights reportedly ship July 27. That date matters more than launch day did.
Each link below shares sources, entities, or timing with this story.
Moonshot AI released Kimi K3, a sparse mixture-of-experts activating 16 of 896 experts per token. That's about 1.8% of the pool live at any moment, with a 1M-token context window and native vision. Two new architectural pieces show up: Kimi Delta Attention and Attention Residu...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
Simon Willison pulled the numbers out of an FT report sourced to "people with knowledge of the matter": Anthropic's annualized revenue reached $65bn in July, up from $47bn in May. Six thousand customers spend $100,000 or more a year. The company told investors it expects a pro...
Moonshot AI dropped Kimi K2.7-Code on Hugging Face on June 12. The specs are loud: 1T-parameter MoE with 32B active across 384 experts, a 256K context window, Modified MIT license, tuned for long-horizon agentic software engineering (MarkTechPost). Moonshot reports +21.8% on K...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
An open-weight Chinese frontier model is now a dropdown option in Microsoft's coding product. That happened before anyone finished characterizing what the model does. GitHub's changelog dated August 6 makes Kimi K3 generally available across Copilot Pro, Pro+, Max, Business an...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.