Fetching from the wire…
Public story · 2026-08-31 · high
The recipe stacks four cost cuts on consumer RTX 5090s and gets close to Qwen2.5-1.5B without a data center.
Why now: Puro-2B's paper posted to arXiv in August 2026, making the training recipe public for the first time.
Puro-2B trains a 2B parameter language model from scratch on consumer RTX 5090 GPUs for under $6,900, per the Puro-2B paper on arXiv. It runs pretraining in FP8 across up to 1.4 trillion tokens and gets close to Qwen2.5-1.5B under the authors' own evaluation protocol.
The authors cite more than $1.5 million to train Llama-3.2-3B and more than $700,000 to reproduce SmolLM3-3B. Puro-2B's budget is more than 200 times below the Llama figure, on hardware a hobbyist can own instead of renting from a cloud provider.
No single trick explains the gap. The savings stack four techniques: RTX 5090s over data-center accelerators, FP8 precision, hyperball optimization, and curriculum model averaging. Cut any one and the budget likely creeps back up, though the paper doesn't say by how much each piece contributes on its own.
What's missing is downstream testing. Coming close to Qwen2.5-1.5B on the authors' chosen benchmarks isn't the same as matching it on benchmarks a skeptical reader would pick. The paper doesn't say how the model behaves on tasks outside that evaluation set.
Each link below shares sources, entities, or timing with this story.
Meta released Llama / Shared entity: GPUs / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Meta released Llama); both cover GPUs; overlapping topics (against, gpus).
Linked by a graph relationship (Meta released Llama); both cover GPUs; overlapping topics (actually, against).
Meta released Llama / Shared entity: Llama / Shared topic / Earlier coverage
Linked by a graph relationship (Meta released Llama); both cover Llama; overlapping topics (against, model).
Meta released Llama / Shared entity: RTX / Shared topic / Earlier coverage
Linked by a graph relationship (Meta released Llama); both cover RTX; overlapping topics (against, model).
Meta released Llama / Shared entity: GPUs / Shared topic / Earlier coverage
Linked by a graph relationship (Meta released Llama); both cover GPUs; overlapping topics (builder, could).
Meta released Llama / Shared topic / Tension
Linked by a graph relationship (Meta released Llama); overlapping topics (against, could, hardware); pushes against this story (against).
Meta released Llama / Same source domain / Shared topic
Linked by a graph relationship (Meta released Llama); reported by the same outlet (arxiv.org); overlapping topics (against, beat).
Meta released Llama / Shared entity: Llama / Earlier coverage
Linked by a graph relationship (Meta released Llama); both cover Llama; earlier Llama coverage from 2026-07-19.