Fetching from the wire…
Models2026-09-21 · source-backed
ciprianveg published benchmarks and runtime patches: about 30 tok/s sustained during heavy code generation peaking near 38, 750-910 tok/s prefill after NCCL topology changes and a dual-switch layout, and 136 t/s peak under concurrency without starving KV cache. Networking is dual MikroTik CRS804-4DDQ switches with 4x400G-to-4x100G breakouts, on a customized gb10-vllm with dspark wrappers and custom MLA/KV kernels. Stable multi-hundred-thousand-token agentic runs with 500k compaction. All patches and build scripts published.
Each link below shares sources, entities, or timing with this story.
An r/LocalLLaMA post at 1,330 upvotes reports the first run of full K3, Moonshot's 2.8T open-weight MoE, on a 16x NVIDIA GB10 cluster with dspark speculative decoding: 20+ tok/s average, 38 peak, 750 prefill. That's roughly $64K of hardware for frontier-adjacent tokens at your...
Hy4-Preview runs 256 routed experts plus one always-active shared expert with top-8 routing, combining Multi-head Latent Attention, DeepSeek Sparse Attention with shared indexer layers, gated MLA with learnable attention sinks, and Independent Hyper-Connections replacing the p...
Moonshot released K3's open weights July 26 with official guidance calling for 64+ accelerators. WASTE (1,366 stars, created July 28) runs it on a 64GB MacBook Pro at 0.45-0.62 tok/s, keeping the 27.28GB trunk resident and streaming experts from NVMe with 3-bit residual vector...
Alongside WASTE, gavamedia/deltafin (603 stars, created July 28) runs full K3 on a single device with an OpenAI-compatible server. Moonshot published K3's open weights July 26-27 at 2.8T parameters; within 48 hours two separate projects appeared whose entire purpose is fitting...
Moonshot exposes an Anthropic-compatible endpoint, so pointing Claude Code at K3 means setting the Anthropic base URL and supplying a Moonshot key. No new CLI, no config rewrite. Hosted at $3/$15 per Mtok, same tier as Claude Sonnet 4.6, and Artificial Analysis scores K3 at 57...
The file is 1.56TB. That's the first thing you notice about the moonshotai/Kimi-K3 Hugging Face repo that went live today. 2.8 trillion total parameters, 104B activated, 896 routed experts with 16 selected plus 2 shared per token, 93 layers split 69 Kimi Delta Attention and 24...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.