Kimi K3's Technical Report Names the Four Tricks Behind a Claimed 2.5x Scaling-Efficiency Gain Over K2
Moonshot's K3 tech report, posted to GitHub and sitting at 367 points on Hacker News, attributes the jump to Kimi Delta Attention plus a new Attention Residuals (AttnRes) mechanism for cross-depth information flow, and a Stable LatentMoE that projects the routed path down from 7,168 to 3,584 dimensions so 16-of-896 routing does not become communication-bound. Training stability at 2.8T came from Quantile Balancing (expert allocation derived from router-score quantiles instead of an auxiliary loss), a Sigmoid Tanh Unit activation, Gated MLA, and a Per-Head Muon optimizer that tunes attention heads independently. Moonshot claims the combination yields ~2.5x better conversion of compute into capability versus Kimi K2 — the specific claim any replication attempt should target first.
↳ Follow the thread