Fetching from the wire…
Research2026-09-10 · source-backed
Dwarkesh Patel and Jerry Han decomposed 2019-2025 pretraining gains at a 1e19 FLOPs budget: a 3.24x gap, or 1.51x per year for data against 1.24x for models. Additive data and model effects explain 88% of performance variance with almost no interaction term. Their caveat is the one to carry: architecture work mostly bought the ability to scale rather than direct efficiency, so the two aren't substitutes and the decomposition shouldn't be read as "architecture doesn't matter."
Each link below shares sources, entities, or timing with this story.
In a June 8 essay, Patel defines intelligence as sample efficiency, argues models have barely improved on it, and says the real gains come from widening data distribution and scaling the compute that manufactures data (RL reframed as verifier-guided synthetic data). His conclu...
Justin Wang and Dan Robinson's RSI Simulator is a browser game where you run an AI lab allocating labor, compute and data toward superintelligence, built on the Elasticity Institute's economics of recursive self-improvement (Paradigm). The model turns on elasticities, chiefly...
Check your GitHub Copilot settings right now. As of April 24, GitHub's updated privacy policy flipped the default for all Copilot Free, Pro, and Pro+ users: your interaction data, including prompts, suggestions, and code snippets from your context, now trains AI models unless...
Most looped-transformer results compare at fixed model size, conflating architecture with extra compute. SMELT matches per-token FLOPs, non-embedding parameters and KV cache against an unlooped baseline, scaling to 54B non-embedding parameters with a separate Chinchilla-style...
Sony Music Publishing and Warner Chappell filed August 28 in the Northern District of California against Anthropic, CEO Dario Amodei and co-founder Benjamin Mann, over what they call a "brazen campaign of illegally torrenting, scraping and downloading copyrighted works on a ma...
arXiv 2608.06196 pits lexical+dense ranking against a graph encoding prerequisites, data flow and ordering across 117 realistic non-echoing queries. The ranker hits top-5 in 73.5% ±8.0 of cases; graph neighbours at matched token budget lose 11.2 points at p=0.0007. The mechani...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.