kimi-k3-mlx Publishes a Full Reverse-Engineering of Kimi K3's Architecture — and Says It Will Not Run on Any Single Mac
PipeNetwork/kimi-k3-mlx (created 2026-07-27, 284 stars) is an MLX port of Kimi K3 that documents the architecture in unusual detail: 896 routed experts at top-16, 2 shared experts per token, 93 layers split 69 Kimi Delta Attention / 24 gated MLA, a novel SiTU-GLU activation, AttnRes residual-stack mixing every 12 layers, and LatentMoE experts running in 3584 dims rather than the 7168-d residual. Routed experts are 97.94% of the 2,780 B parameters, and the accounting validates against the published 1.561 TB repo size exactly. The README opens with a warning that flatly contradicts the same week's single-machine claims: the smallest tier is ~870 GB against a 512 GB ceiling on the largest Apple Silicon machine that exists.
Source
↳ Follow the thread