Fetching from the wire…
Models2026-07-30 · source-backed
The partnership put the 2.8-trillion-parameter K3 on Together's serverless inference API on July 27, the same day Moonshot released full weights. K3 ships native vision and a 1M-token context window and is the first open-weight model in the 3T-parameter class. Developers get the full Moonshot lineup behind one API with fine-tuning on their own data, no new SDK per release.
Each link below shares sources, entities, or timing with this story.
The open-weight race just changed constraint. Moonshot AI suspended all new consumer subscriptions on July 20, roughly 48 hours after Kimi K3 launched, because request volume pushed its compute cluster to capacity. Remaining GPUs are reserved for existing paid subscribers. Tec...
Moonshot released it July 16: 2.8T parameters, MoE routing 896 experts with 16 active per token, native multimodal input, 1M-token context, with full open weights promised by July 27. Model overview here. It ships MXFP4 (4-bit float with per-block scaling) from day one, puttin...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
paddo.dev makes the most contrarian read: the letter's substance isn't openness but paragraph nine, defending distillation as "a widely used technique for model improvement" and urging policymakers against "conflating legitimate model development techniques with misappropriati...
OpenAI posted first benchmark results for Jalapeño, its Broadcom co-developed inference ASIC, claiming 1.5x to 1.9x more throughput per kilowatt and 1.7x to 3.6x lower end-to-end latency than Nvidia GB200 and GB300 rack systems, measured on the SemiAnalysis InferenceX suite. T...
Moonshot AI released Kimi K3, a sparse mixture-of-experts activating 16 of 896 experts per token. That's about 1.8% of the pool live at any moment, with a 1M-token context window and native vision. Two new architectural pieces show up: Kimi Delta Attention and Attention Residu...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.