First Reported Run of Full Kimi K3 on a 16x GB10 Cluster: 20+ tok/s Average, 38 tok/s Peak, 750 tok/s Prefill
An r/LocalLLaMA post at 1,330 upvotes / 266 comments reports what the author calls the first run of full Kimi K3 — Moonshot's 2.8-trillion-parameter open-weight MoE — on a 16x NVIDIA GB10 (DGX Spark) cluster with dspark speculative decoding, hitting 20+ tokens/sec average on llama-benchy's coherent corpus, 38 tok/s peak, and 750 tok/s prefill. K3's weights went public July 26; the open question since then has been what it actually takes to self-host 2.8T parameters, and the NVIDIA developer forums show a parallel thread working the same problem on 24 Sparks (96 KDA heads, 93 layers, 896 routed experts, TP24 vs TP8/PP3 tradeoffs). Sixteen Sparks is roughly $64K of hardware for frontier-adjacent tokens at your desk — the number that makes the 'largest open-weight model ever' claim operationally real rather than theoretical.
↳ Follow the thread