Reddit
A Hosting Team Is Putting Kimi K3 on A100s to See If It Holds Up — and Says the A100 Math Is 'Already Rough'
An r/LocalLLaMA operator thread (505 upvotes, 134 comments) documents a live plan to serve K3 on A100s, H200s and B300s, with A100/H200 results promised this week and a B300 cluster going up over the weekend. The 'rough math' framing tracks with vLLM's own guidance published the same day: 8x B300 or 16x B200/GB200 is the practical floor, and the ~594GB MXFP4 weight file plus 1M-token KV pressure leaves little headroom on 80GB Ampere parts without heavy expert parallelism over the network. This is the most useful kind of post for anyone budgeting inference — a team publishing the failure case before the vendor benchmark.
↳ Follow the thread