Fetching from the wire…
Infra2026-09-16 · source-backed
PR #57050 fixes get_num_blocks_to_allocate miscalculating physical blocks for long requests loading prefix cache from external KV storage such as Mooncake, so the request is refused admission and sits in the waiting queue forever, stalling the whole system. Maintainers trace it to a likely regression from #53614. The reproduction needs 8 B300s serving Kimi-K3 with specific flags, narrow enough that most operators would have experienced it as an unexplained hang.
Each link below shares sources, entities, or timing with this story.
The minute Fable 5 and Mythos 5 went dark for foreign nationals, r/LocalLLaMA found its answer. Moonshot AI's Kimi K2.7 Code is a 1T-parameter MoE (32B active, 384 experts), 256K context, shipped under a Modified MIT license. The headline number that's getting it pulled: 81.1...
Cursor shipped Composer 2 on March 19, marketing it as a proprietary in-house model. Within 24 hours, a developer found the API routing to kimi-k2p5-rl-0317-s515-fast. The model powering the most-hyped coding tool update of the month was Moonshot AI's Kimi K2.5 with continued...
The file is 1.56TB. That's the first thing you notice about the moonshotai/Kimi-K3 Hugging Face repo that went live today. 2.8 trillion total parameters, 104B activated, 896 routed experts with 16 selected plus 2 shared per token, 93 layers split 69 Kimi Delta Attention and 24...
agentic-kv-cache simulates cross-request prefix caching against real traces, not synthetic ones: 68,266 requests across 393 Claude Code sessions at 64-token blocks, plus Mooncake traces at 512-token blocks. It models prefix-contiguous hits, radix eviction constraints and pinne...
Moonshot AI replaces standard fixed residual connections with softmax attention over preceding layer outputs. Already in production at 48B scale (Kimi Linear). Consistent scaling improvement validated across model sizes. 1,330 HuggingFace upvotes. Source
Moonshot AI released Kimi K3, a sparse mixture-of-experts activating 16 of 896 experts per token. That's about 1.8% of the pool live at any moment, with a 1M-token context window and native vision. Two new architectural pieces show up: Kimi Delta Attention and Attention Residu...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.