Fetching from the wire…
Infra2026-09-14 · source-backed
PR #55737 reuses the vllm._flashkda_C kernel already built for Kimi-K3, because FlashKDA implements the same bounded-gate recurrence and drops straight in. Per-layer on GB300: 8x2048 tokens from 183us to 49us, 4x8192 from 555us to 146us. End-to-end on 4x GB300 TP4 with prefix caching off, mean TTFT fell 12.2% at 8x2048. Selection is automatic on SM90/SM10x/SM12x with bf16 and head_dim 128, and additional_config.kda_prefill_backend=triton keeps the old path.
Each link below shares sources, entities, or timing with this story.
OSTP Director Michael Kratsios posted July 22 that Moonshot built "a sophisticated internal platform to conduct large scale distillation against U.S. models," switching access methods to avoid detection, and acquired GB300-equipped servers plus GB300 access in Thailand. TechCr...
The assumption that proprietary models own the coding benchmark crown just broke. Moonshot AI's Kimi K2.6 leads on 5 of 8 major agentic coding benchmarks while being the only open-weight model in the top tier. SWE-Bench Pro: 58.6% vs GPT-5.4's 57.7% and Claude Opus 4.6's 53.4%...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
The open-weight race just changed constraint. Moonshot AI suspended all new consumer subscriptions on July 20, roughly 48 hours after Kimi K3 launched, because request volume pushed its compute cluster to capacity. Remaining GPUs are reserved for existing paid subscribers. Tec...
The minute Fable 5 and Mythos 5 went dark for foreign nationals, r/LocalLLaMA found its answer. Moonshot AI's Kimi K2.7 Code is a 1T-parameter MoE (32B active, 384 experts), 256K context, shipped under a Modified MIT license. The headline number that's getting it pulled: 81.1...
Moonshot AI dropped Kimi K2.7-Code on Hugging Face on June 12. The specs are loud: 1T-parameter MoE with 32B active across 384 experts, a 256K context window, Modified MIT license, tuned for long-horizon agentic software engineering (MarkTechPost). Moonshot reports +21.8% on K...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.