Fetching from the wire…
Research2026-09-03 · source-backed
Interpretability researchers have assumed shared structure between knowledge domains makes unlearning harder; nobody had tested it. The authors trained six 254M-parameter models on English Wikipedia with graded disentanglement between biology and non-biology knowledge, then applied three unlearning methods to each. At fixed forgetting, the most disentangled models incurred roughly 4x lower retain cost under two methods and 1.3x under the third. The intervention changes only the model, not the data or the algorithm, so this is causal evidence rather than correlation.
Each link below shares sources, entities, or timing with this story.
Piotr Wilam crossed Python and Rust with Qwen2.5-Coder-7B and DeepSeek-Coder-V1-6.7B, inventorying grammatical concepts (58 Python, 57 Rust) identically in all four cells. Which concepts earn dedicated circuitry is set by the task — the models agree at Spearman rho = 0.638 for...
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
Logistic-regression probes on a coding agent's hidden states can decode whether code will parse and pass tests at AUC up to 0.83, and those internal representations run ahead of the agent's own edits, predicting outcomes as much as 25 steps in advance (arXiv). The authors call...
DSEffi-Bench covers 1,000 instances across 10+ libraries with stress-testing harnesses and human-validated references, evaluated on 16 models (arXiv 2608.30248). GPT-5.4 leads correctness at 66.9% Pass but its 71.7% efficiency score barely beats GPT-5.4-mini's 71.6% despite so...
Puro-2B trains from scratch on up to 1.4 trillion tokens in FP8 on consumer GPUs, approaching Qwen2.5-1.5B under the authors' protocol, against a stated $1.5M+ to train Llama-3.2-3B and $700K+ to reproduce SmolLM3-3B. The savings stack rather than coming from one trick: hardwa...
A fleet evaluation across 46 endpoints from six vendors found a recognition-enforcement gap: source-format features are linearly decodable from activations and models verbally identify forged authority when asked, but some configurations still emit the conflicting tool call. A...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.