r/LocalLLaMA Does the Kimi K3 Memory Math Before the Weights Land: Only 8x B300 Fits in a Single Node
A deployment team posted worked memory arithmetic ahead of the July 27 drop (494 upvotes, 134 comments) and the numbers are brutal for older silicon: 8x A100 80GB gives 640GB against ~1.4TB of weights, meaning three nodes before any KV cache, and Ampere has no FP4 or FP8 tensor cores so you are dequantizing or running INT4 kernels the release never targeted. 8x H200 reaches ~1.13TB — still a two-node minimum with interconnect cost on every token. Only 8x B300 (~2.3TB) fits the whole model in one node with headroom for long-context KV cache, and Blackwell's native FP4 is clearly what Moonshot quantized for. They also note Moonshot's model card is unusually candid about weaknesses: quality drops if your agent harness truncates thinking history, it tends to act rather than ask when instructions are ambiguous, and chat experience still trails Fable 5 and Sol even where benchmarks are close.
↳ Follow the thread