Fetching from the wire…
Public story · 2026-09-18 · high
The effect held only when models saw the rule demonstrated in training, tested up to 110 billion parameters and 1 billion midtraining tokens.
Why now: The paper posted in September, testing midtraining against competing finetuning data at up to 110 billion parameters.
A small dose of finetuning data wiped out a model's midtraining-installed motivation, according to a paper posted to arXiv.
Researchers ran the test at up to 110 billion parameters and 1 billion midtraining tokens. That scale matters for anyone treating midtraining as a way to install durable values before finetuning starts.
The paper continued pretraining on large volumes of alignment-relevant documents. It then tested whether midtraining could tip a model's motivation when post-training data left it ambiguous between two options.
Midtraining worked. It tipped the balance. Then a small slice of finetuning data suggesting a competing motivation wiped the effect out entirely.
The paper's second finding cuts deeper. Rules were only reliably learned when a demonstration of the rule appeared in midtraining or post-training data. A rule stated but never demonstrated didn't generalize. That undercuts the idea that midtraining teaches behavior a model never sees performed.
If a handful of finetuning examples can undo midtraining, midtraining was never installing a value. It was setting a prior that finetuning overwrites on contact. Watch whether labs leaning on midtraining as an alignment layer test it against finetuning data built to disagree with it. This paper suggests a handful of contradicting examples is enough to undo it.
Each link below shares sources, entities, or timing with this story.
A measurement paper published August 28 ran one adaptive adversary against a seven-layer stack and found failure correlation positive in all fifteen measurable pairs, phi between 0.30 and 0.75, with the joint residual exceeding the multiplicative prediction by up to 0.172. The...
Danish Foundation Models trained it from scratch on 161 datasets. Across 20 benchmarks spanning English, math and code, and Danish, it beats the original HRM-Text 1B, sets a new Danish state of the art, and competes with Qwen 3.5 4B and Gemma 4 E2B. Weights are on Hugging Face...
arXiv 2608.06337 settles an open question on the monotone-adversary model, where an adversary appends examples all labeled correctly by the target hypothesis but chosen after seeing the clean sample. The extra logarithmic factor is inherent, not algorithmic: minimax expected e...
arXiv 2607.26935 argues the human-vs-bot label space can't represent agent traffic: an MLP binary classifier misroutes 39.1% of real agent sessions as human, a SAINT transformer 34.5%, while adding an explicit third class yields agent F1 = 1.000 across all 30 runs. Against a f...
This arXiv work claims training only one layer during RL post-training matches full-parameter RL fine-tuning, and it was surfacing on Hacker News. If it holds, RLHF/RLVR post-training gets dramatically cheaper, because you're touching a fraction of the network. Big "if." But t...
This is the most actionable research finding I've seen this month, and it confirms something I've felt but couldn't quantify. Paper arXiv:2604.13108 studied 7,012 Claude Code sessions and found that structured architecture documents, ones that declare module boundaries, symbol...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.