Sources
NCP-ArchPreview reaches OLMo-3-7B's final pretraining loss on 51.3% of the tokens by predicting multi-token concepts
arXiv 2609.10715 (2026-09-09, top Hugging Face daily paper on 09-11 with 109 upvotes) trains an 8.9B model on 5.73T Dolma-3 tokens. Alongside next-token prediction, it learns Next Concept Prediction over a product-quantized concept vocabulary built from its own hidden states. After full pretraining it beats OLMo-3-7B by 2.45 points on the downstream macro-average and by 5.99 on GSM8K. The authors call it the largest latent-space language model shown so far. It adds evidence alongside this month's looped-transformer work that architecture changes, not only more tokens, still improve pretraining efficiency.
↳ Follow the thread