Research
On-Policy Distillation Barely Distills: A Fixed Negative Advantage Matches the Teacher
Quantifying teacher supervision during on-policy distillation shows substantial noise that gets worse as the teacher scales, yet the student converges to comparable performance whether the noisy supervision is kept or stripped out. Learning concentrates on low log-probability tokens, and swapping teacher-provided advantages for a single fixed negative advantage matches full OPD, implying the method works mostly by suppressing tail tokens and needs no teacher at all. The resulting supervision-free method, OPSA, uses entropy-adaptive negative advantages and improves Avg@32 on AIME24 by 35.41 points over base Qwen3-1.7B, a 263% relative gain, and beats OPD itself by 16.77 points.
↳ Follow the thread