Fetching from the wire…
Models2026-07-18 · source-backed
Moonshot released it July 16: 2.8T parameters, MoE routing 896 experts with 16 active per token, native multimodal input, 1M-token context, with full open weights promised by July 27. Model overview here. It ships MXFP4 (4-bit float with per-block scaling) from day one, putting the full weight footprint around 1.4 TB rather than the ~5.6 TB bf16 would need. API pricing is $3.00/M input on cache miss, $0.30 on hit, $15.00/M output. The quantization choice is what makes self-hosting even theoretically discussable at this scale, and I'd bet that was the point.
Each link below shares sources, entities, or timing with this story.
The open-weight race just changed constraint. Moonshot AI suspended all new consumer subscriptions on July 20, roughly 48 hours after Kimi K3 launched, because request volume pushed its compute cluster to capacity. Remaining GPUs are reserved for existing paid subscribers. Tec...
Moonshot released K3's open weights July 26 with official guidance calling for 64+ accelerators. WASTE (1,366 stars, created July 28) runs it on a 64GB MacBook Pro at 0.45-0.62 tok/s, keeping the 27.28GB trunk resident and streaming experts from NVMe with 3-bit residual vector...
The file is 1.56TB. That's the first thing you notice about the moonshotai/Kimi-K3 Hugging Face repo that went live today. 2.8 trillion total parameters, 104B activated, 896 routed experts with 16 selected plus 2 shared per token, 93 layers split 69 Kimi Delta Attention and 24...
An open-weight Chinese frontier model is now a dropdown option in Microsoft's coding product. That happened before anyone finished characterizing what the model does. GitHub's changelog dated August 6 makes Kimi K3 generally available across Copilot Pro, Pro+, Max, Business an...
Alongside WASTE, gavamedia/deltafin (603 stars, created July 28) runs full K3 on a single device with an OpenAI-compatible server. Moonshot published K3's open weights July 26-27 at 2.8T parameters; within 48 hours two separate projects appeared whose entire purpose is fitting...
PR #26062, "server: support MCP stdio," by ngxson, merged into ggml-org/llama.cpp on July 25 (r/LocalLLaMA). It landed alongside #26061 (vendored subprocess.h, merged July 24) and pwilkin's #26075 integration-and-tests PR. Until now, llama-server's web UI could only talk to MC...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.