Reddit
Kimi K3's full 2.8T parameters run on a 16x GB10 cluster at 30 t/s of coding throughput
ciprianveg published benchmarks and runtime patches for running the full moonshotai/Kimi-K3 across 16 GB10 nodes: ~30 tok/s sustained during heavy code generation (peaking near 38), 750-910 tok/s prefill after NCCL topology changes and a dual-switch layout, and 136 t/s peak under concurrency without starving KV cache. Networking is dual MikroTik CRS804-4DDQ switches with 4x400G-to-4x100G breakouts; the stack is a customized gb10-vllm with dspark wrappers and custom MLA/KV kernels, and he reports stable multi-hundred-thousand-token agentic runs with 500k compaction. All patches and build scripts are published.
↳ Follow the thread