Fetching from the wire…
Research2026-09-07 · source-backed
1-shot on-policy distillation was consistently effective across every sampled training example, with harder examples yielding larger gains (arXiv 2609.05198). The analysis attributes the improvement to the longer CoT paths hard problems naturally generate, which keep the student aligned with the teacher over long horizons and teach reflection patterns absent from short CoTs, not to high token entropy. A hard-example-only selection method, including "unsolvable" examples exceeding the teacher's own capability, trained successfully with 8 selected examples across four models from 1.5B to 7B.
Each link below shares sources, entities, or timing with this story.
FACE-Eval varies where a preference cue is delivered, user message or tool return, across 5,100 samples and 15 open-weight models from 4B to 1.60T parameters (arXiv 2608.29464). Every single model showed lower verbalized commitment for tool-return cues, and unverbalized adopti...
Steering interventions treat a model's recognition that it's being tested as one quantity to suppress. In chain-of-thought, verbalized eval-awareness separates into capabilities-flavored ("testing my ability to follow instructions") and safety-flavored ("testing my boundaries"...
The training-free method builds memories from historical traces summarizing reasoning patterns, key constraints and critical operations, then retrieves them as prefill-side scaffolds. Gains of 21.4, 28.0, 29.5 and 6.61 points on GSM8K, MATH, BBH and MMLU-Sci, with a 1.14-1.49x...
arXiv 2608.04735 points out that monitorability evals overwhelmingly use *explicit* influence, where the prompt tells the model to hide a side task, and monitors catch 60-94% across seven frontier extended-thinking models. Swap in subtle contextual bias and detection drops 41-...
Rewriting an agent's reasoning to read as good-faith engineering while copying every command and output verbatim, so the exploit itself is unchanged, defeats a held-out CoT monitor almost completely. The sharper point for anyone running a monitor in production: headline accura...
arXiv 2607.29529 emits an auditable trace alongside the program using a contract-annotated task graph binding stable responsibility identities to each commitment, implementation, provenance, validation evidence and intervention history. On validation failure a conservative loc...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.