Kimi K3 at full precision runs at 1.00 tokens/sec on a MacBook Pro by streaming 1.45 TB of experts from four SSDs
Hacker News·high signal
argonautlabsai/deltafin published benchmarks running the unquantized 2.8T-parameter Kimi K3 MoE on an M5 Max MacBook Pro with 128 GB RAM, loading expert weights on demand from four SSDs totaling 1.45 TB. Measured decode was 1.00 tok/s over 512 tokens, 1.13 tok/s over 128 tokens, and 0.96 tok/s on a 17-token public prompt, against 0.29 tok/s from the upstream gavamedia/deltafin on an M1 Max. The repo was created 2026-09-08 and hit 273 points on Hacker News the same day at 72 stars, so the interest is in the storage-streaming technique, not the project.