Fetching from the wire…
Research2026-09-01 · source-backed
Measuring teacher supervision during on-policy distillation shows substantial noise that worsens as the teacher scales, yet the student converges comparably whether that supervision is kept or stripped (arXiv 2608.31046). Learning concentrates on low log-probability tokens, and swapping teacher advantages for a single fixed negative advantage matches full OPD. The supervision-free method that follows, OPSA, improves Avg@32 on AIME24 by 35.41 points over base Qwen3-1.7B and beats OPD itself by 16.77. If the teacher is not doing the work, a lot of distillation budget is being spent on nothing.
Each link below shares sources, entities, or timing with this story.
arXiv 2608.03999 holds the Qwen3.5 backbone (0.8B–27B), data, budget, and decoding fixed and swaps only the representation across seven tokenizations. Scaling the backbone 34x barely moves Frechet Music Distance; switching representation halves it. Their PMT stream (10ms timin...
arXiv 2608.24658 measures that prior parallel-reasoning work chased subtask parallelism while the larger share is trial parallelism, where multiple speculative attempts explore, verify and aggregate competing hypotheses simultaneously. It accounts for 65.5% of DeepSeek-V4's pa...
SimpleOPD tackles the practical blocker in on-policy distillation, that teacher and student rarely share a tokenizer, by distilling in shared text space (arXiv 2608.14277). It adds a student reference KL loss and masks the advantages of special termination tokens to stop runaw...
Measuring expert time on two datacenter GPU generations shows it is linear in neither token count (EPLB, LPLB, UltraEP) nor activated-expert count (METRO). Below roughly 156 to 168 tokens, HBM weight streaming dominates so cost attaches to activated replicas. Above it, grouped...
The Agent Lightning result (arXiv 2608.17528) is a 9B model gaining 14.6 points from 3,500 lines of training code and modest compute. The point isn't that a 9B beats anything, it's the cost curve: 6K examples is a dataset a small team can build. Task-specific agentic RL on an...
Visual on-policy distillation normally needs a stronger teacher or privileged supervision (arXiv 2608.14144). S²VOPD inverts the asymmetry: the same model acts as teacher on the original image and student on a strongly augmented view, so the signal comes free of annotations, r...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.