Fetching from the wire…
Models2026-07-29 · source-backed
Sebastian Raschka's July 28 teardown argues K3 is less exotic than the release framing suggests: a scaled production version of Kimi Linear with Kimi Delta Attention as the hybrid attention layer and LatentMoE compressing large linear layers by down-projection. The genuinely new piece is replacing every RoPE layer with NoPE, the first frontier-scale no-positional-embeddings implementation, alongside attention residuals that weight cross-layer residual connections by attention scores for roughly 4% added training cost and 2% added inference cost. He groups K3 with Nemotron 3 Ultra and DeepSeek V4 as a clear industry pivot toward inference efficiency.
Each link below shares sources, entities, or timing with this story.
PipeNetwork/kimi-k3-mlx (created July 27, 284 stars) documents the architecture in unusual detail: 896 routed experts at top-16, 2 shared experts per token, 93 layers split 69 Kimi Delta Attention / 24 gated MLA, a novel SiTU-GLU activation, AttnRes residual-stack mixing every...
Moonshot AI replaces standard fixed residual connections with softmax attention over preceding layer outputs. Already in production at 48B scale (Kimi Linear). Consistent scaling improvement validated across model sizes. 1,330 HuggingFace upvotes. Source
Moonshot AI released Kimi K3, a sparse mixture-of-experts activating 16 of 896 experts per token. That's about 1.8% of the pool live at any moment, with a 1M-token context window and native vision. Two new architectural pieces show up: Kimi Delta Attention and Attention Residu...
The file is 1.56TB. That's the first thing you notice about the moonshotai/Kimi-K3 Hugging Face repo that went live today. 2.8 trillion total parameters, 104B activated, 896 routed experts with 16 selected plus 2 shared per token, 93 layers split 69 Kimi Delta Attention and 24...
The UK AI Security Institute and US CAISI published a joint preliminary cyber evaluation of Moonshot's Kimi K3 (open weights due July 27). On ExploitBench, a Carnegie Mellon benchmark covering 41 post-2023 Chrome V8 vulnerabilities, K3 hit 32% versus 76% for the most cyber-cap...
The changelog has Kimi K3 as a selectable model (Aug 6) and effort levels reaching GA (Aug 7), plus an ROI section in the impact dashboard. Effort levels is the operationally relevant one: review depth as a per-invocation dial rather than a fixed cost, the same knob Claude Cod...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.