Fetching from the wire…
Research2026-09-13 · source-backed
arXiv 2609.11699 starts from the finding that On-Policy Self-Distillation degrades hard-reasoning performance, because conditioning on ground-truth solutions produces artificially confident traces that suppress hedging and penalize the exploratory self-correction hard problems need. NSD inverts it: the model generates a question-specific negative condition, such as acting as a careless reasoner, and the student distribution is pushed away from that self-generated negative teacher. No ground-truth answers and no external supervision, which is what makes it interesting for domains where you have problems and no verified solutions.
Each link below shares sources, entities, or timing with this story.
SkillJack found safety detection on poisoned trajectories ran 98.5% but fell to 11.4% on skills extracted from those same trajectories, with 80% of skill-mediated attacks persisting after the original records were deleted. Distillation launders intent. Cleaning your trace stor...
Uses conditional mutual information to measure distillation-relevant info leakage through API outputs. Learns a transformation that strips distillation info while preserving task accuracy. Timely given Anthropic's industrial-scale distillation disclosure. (arXiv)
arXiv 2609.10263 separates what a persistent agent stores from what it uses, because a superseded fact misleads a current-state answer while remaining necessary for a historical query. A retained archive holds everything; a query-conditioned view governs influence, with same-s...
Diverse Hypothesis Deliberation caches five independently generated messages per problem, then hides and reveals each to the same downstream integrator to measure marginal contribution (arXiv 2608.14375). Across five math and science benchmarks and two model families, wrong-bu...
The Wiggle Framework stress-tested 9 frontier models across 14 judging tasks. The damning part: flips were almost always net-corrupting relative to ground truth. Pressure moved judges away from the right answer, not toward it. arXiv 2608.12645 If you use LLM-as-judge anywhere...
Scoring each retrieved chunk and dropping failures assumes one chunk is a sufficient premise; multi-hop questions are built so none is. Entailment scoring reaches 0.643/0.523/0.560 AUC on HotpotQA, 2Wiki, and MuSiQue against 0.951 on single-hop SQuAD, and per-chunk gating was...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.