Fetching from the wire…
Public story · 2026-07-30 · high
A reverse-engineered parameter count puts Kimi K3's smallest deployable tier near 870GB, past the 512GB ceiling on the biggest Apple Silicon machine sold.
Why now: The repo went up July 27, directly countering single-Mac claims about Kimi K3 that were circulating the same week.
Kimi K3's full architecture doesn't fit on a single Mac, according to a reverse-engineered parameter map published on GitHub July 27. The repo, from PipeNetwork, has pulled in 284 stars for laying out exactly what's inside the model.
That matters for anyone budgeting hardware to run frontier open-weight models locally. Routed experts alone account for 97.94% of Kimi K3's 2,780B total parameters, per the repo's breakdown. That accounting validates exactly against the model's published 1.561TB repo size.
The detail is granular. 896 routed experts get selected at top-16 per token, plus 2 shared experts. They spread across 93 layers split 69 Kimi Delta Attention to 24 gated MLA. A SiTU-GLU activation function and an AttnRes mechanism that mixes the residual stream every 12 layers aren't things I've seen documented elsewhere. The LatentMoE experts run in a compressed 3584-dimension space, half the model's 7168-d residual stream.
The README doesn't hedge on this point. The smallest deployable tier needs roughly 870GB, and the largest Apple Silicon machine on the market tops out at 512GB. No single Mac clears that gap.
Kimi K3 was built as a cluster model, not a local one. Any claim of running the full weights on a single Mac is quantized, distilled, or false. Check the math before you believe one.
The repo went up July 27, directly countering single-Mac claims about Kimi K3 that were circulating the same week.
Each link below shares sources, entities, or timing with this story.
The file is 1.56TB. That's the first thing you notice about the moonshotai/Kimi-K3 Hugging Face repo that went live today. 2.8 trillion total parameters, 104B activated, 896 routed experts with 16 selected plus 2 shared per token, 93 layers split 69 Kimi Delta Attention and 24...
Sebastian Raschka's July 28 teardown argues K3 is less exotic than the release framing suggests: a scaled production version of Kimi Linear with Kimi Delta Attention as the hybrid attention layer and LatentMoE compressing large linear layers by down-projection. The genuinely n...
Moonshot released K3's open weights July 26 with official guidance calling for 64+ accelerators. WASTE (1,366 stars, created July 28) runs it on a 64GB MacBook Pro at 0.45-0.62 tok/s, keeping the 27.28GB trunk resident and streaming experts from NVMe with 3-bit residual vector...
Moonshot AI released Kimi K3, a sparse mixture-of-experts activating 16 of 896 experts per token. That's about 1.8% of the pool live at any moment, with a 1M-token context window and native vision. Two new architectural pieces show up: Kimi Delta Attention and Attention Residu...
Alongside WASTE, gavamedia/deltafin (603 stars, created July 28) runs full K3 on a single device with an OpenAI-compatible server. Moonshot published K3's open weights July 26-27 at 2.8T parameters; within 48 hours two separate projects appeared whose entire purpose is fitting...
An open-weight Chinese frontier model is now a dropdown option in Microsoft's coding product. That happened before anyone finished characterizing what the model does. GitHub's changelog dated August 6 makes Kimi K3 generally available across Copilot Pro, Pro+, Max, Business an...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.