Sources
Puro-2B pretrains a 2B model on consumer RTX 5090s for under $6.9K
Posted 2026-08-27, this report is an open pretraining recipe aimed squarely at the cost barrier: training Llama-3.2-3B runs over $1.5M and reproducing SmolLM3-3B needs over $700K. Puro-2B trains from scratch on up to 1.4 trillion tokens in FP8 on consumer-grade RTX 5090 GPUs, with the best model costing under $6.9K in compute and approaching Qwen2.5-1.5B under the authors' evaluation protocol. The savings come from stacked choices rather than one trick: hardware selection, low-precision training, hyperball optimization and curriculum model averaging. This is the recipe an independent builder can actually run.
↳ Follow the thread