Sebastian Raschka's Kimi K3 Teardown: It Is Kimi Linear Scaled From 48B to 2.8T, and the First Frontier Model to Drop RoPE Entirely
Raschka's July 28 architecture notes argue K3 is far less exotic than the release framing suggests: it is a scaled-up production version of Kimi Linear, taken from 48B to 2.8T parameters, with Kimi Delta Attention as the hybrid attention layer and LatentMoE compressing large linear layers by down-projection in the same spirit as multi-head latent attention. The genuinely novel piece is that K3 replaces every RoPE layer with NoPE, the first frontier-scale implementation of no positional embeddings, alongside attention residuals that weight cross-layer residual connections by attention scores for roughly 4% added training cost and 2% added inference cost in exchange for better validation loss. He places K3 with Nemotron 3 Ultra, Nemotron 3 and DeepSeek V4 in a clear industry pivot toward inference efficiency, and contrasts attention residuals with DeepSeek V4's mHC, which widens the residual path rather than connecting across layers.
↳ Follow the thread