Reddit
Someone Pre-Trained a 1.02B Kimi-K3 Replica on 5 Billion Tokens for $250 and Beat GPT-2 124M on HellaSwag
A builder rented GPUs on Modal and pre-trained a 1.02B-parameter model (145M active per token) on exactly 5,000,003,584 decontaminated tokens for $250, keeping K3's actual architecture including Kimi Delta Attention, Gated MLA, attention residuals, LatentMoE with the aux-loss-free balancer, and K3's unmodified 163,840-token tokenizer. It scores 33.4% on HellaSwag against GPT-2 124M's 28%, with no instruction tuning at all. The model is roughly one two-thousandth of K3 by size, and the full tutorial plus a companion guide to hosting the real 2.8T K3 on eight B300s is published on books.vizuara.ai.
↳ Follow the thread