Fetching from the wire…
Research2026-09-04 · source-backed
Prior work fuses OPD's dense token-level supervision with RLVR's sparse reward in a single step, either weighted-additive or teacher-modulated advantage rescaling. A plain two-stage OPD-then-RL scheme beats pure OPD, pure RLVR and all joint baselines across logic and math benchmarks. The mechanism: OPD expands the student's coverage of teacher-supported solutions and RL sharpens within that support, while joint optimization makes the signals interfere. The OPD validation score tells you when to switch, and OPD is a better cold start for RL than SFT. arXiv 2609.04108
Each link below shares sources, entities, or timing with this story.
Measuring teacher supervision during on-policy distillation shows substantial noise that worsens as the teacher scales, yet the student converges comparably whether that supervision is kept or stripped (arXiv 2608.31046). Learning concentrates on low log-probability tokens, an...
MoRe learns a codebook of steering vectors, each encoding a latent role, then uses a query-aware router to fuse them into a single composed vector for single-turn inference. The backbone stays frozen; training is a three-stage SFT curriculum plus GRPO. Across reasoning and per...
arXiv 2607.23740 introduces a psychology-grounded benchmark of 3 primary dimensions, 17 secondary, 71 task paradigms, controlling question format, narrative perspective, and context length across 284 shared scenarios and 3,481 expert-verified instances. Not one of the 17 secon...
arXiv 2607.23982 adapts Holmström's team moral-hazard model into a game where an agent can keep an immediate local reward or pay a query cost to surface a hidden safety fact that mainly helps another agent's downstream decision. Base behavior splits into two failure modes: pre...
This arXiv work claims training only one layer during RL post-training matches full-parameter RL fine-tuning, and it was surfacing on Hacker News. If it holds, RLHF/RLVR post-training gets dramatically cheaper, because you're touching a fraction of the network. Big "if." But t...
Diverse Hypothesis Deliberation caches five independently generated messages per problem, then hides and reveals each to the same downstream integrator to measure marginal contribution (arXiv 2608.14375). Across five math and science benchmarks and two model families, wrong-bu...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.