Fetching from the wire…
Public story · 2026-09-11 · high
The model predicts multi-token concepts alongside individual words and beats a same-size baseline on math and reasoning after training on 5.73 trillion tokens.
Why now: As of September 11, 2026, the efficiency claim hasn't been checked by anyone outside the paper's own team.
NCP-ArchPreview reached OLMo-3-7B's final pretraining loss using 51.3% of the training tokens, according to the NCP-ArchPreview paper posted to arXiv. Training runs at this scale cost real money in GPU hours, so cutting the token count nearly in half while hitting the same loss matters, if the result holds outside one lab's setup.
The method trains an 8.9B parameter model on 5.73 trillion Dolma-3 tokens. Alongside standard next-token prediction, it also predicts the next "concept." That's a multi-token unit pulled from a vocabulary built by quantizing the model's own hidden states.
After full pretraining, NCP-ArchPreview also outscored OLMo-3-7B on benchmarks. It beat the baseline by 2.45 points on a downstream macro-average and by 5.99 points on GSM8K, a grade-school math benchmark for multi-step reasoning. The authors call it the largest latent-space language model demonstrated so far.
The paper doesn't settle whether the concept-prediction objective itself does the work, or whether the result depends on the specific Dolma-3 mix or the 8.9B scale. One lab's numbers from a single training run aren't a proven technique yet. Efficiency claims this size tend to shrink once someone else reproduces them at a different scale or with a different tokenizer.
Each link below shares sources, entities, or timing with this story.
The pattern most agent memory systems use, dumping a retrieved trajectory into context, degrades as traces lengthen and as source-specific values diverge from the target. QCR replaces the dump with a structured memory holding reusable procedures plus bindings, applicability co...
Gemini 3.6 Flash launched July 21 at $1.50/1M in, $7.50/1M out, claiming 17% fewer output tokens than 3.5 Flash, DeepSWE code precision up from 37% to 49%, OSWorld-Verified computer use at 83% (from 78.4%), and a knowledge cutoff finally moved from January 2025 to March 2026....
Puro-2B trains from scratch on up to 1.4 trillion tokens in FP8 on consumer GPUs, approaching Qwen2.5-1.5B under the authors' protocol, against a stated $1.5M+ to train Llama-3.2-3B and $700K+ to reproduce SmolLM3-3B. The savings stack rather than coming from one trick: hardwa...
62,876 stars, ~313/day. Compresses tool outputs, logs, RAG chunks and files before they hit the model, and unusually publishes the other side of the trade: GSM8K holds at 0.870 (±0.000), TruthfulQA *improves* 0.530 → 0.560, SQuAD v2 and BFCL retain 97%. Real workloads: 92% sav...
The first systematic measurement of PyPI import cost covers the 500 most-downloaded packages sampled quarterly over five years, under CPython 3.9 through 3.14, on Apple M5/macOS and Intel Xeon/Linux. Half of packages import in under 6 ms but p99 is 354 ms. First import after i...
book-to-skill (14,194 stars, +595 today) compiles a PDF into a ~4K-token SKILL.md plus ~1K-token per-chapter files loaded on demand, reporting 24-51x fewer tokens than dumping the book and ~5K tokens resident versus 119K-256K for a context dump, at roughly $1 per book to conve...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.