Relay-OPD Fixes On-Policy Distillation's Prefix Failure by Letting the Teacher Briefly Take the Wheel, Cutting Training Trajectories in Half
An eight-author paper (arXiv 2607.26057, July 28) identifies prefix failure in on-policy distillation: once a student commits to a wrong reasoning direction, every subsequent token builds on the deviation, producing misdirected continuations that elicit unreliable supervision and burn compute. The authors observe a teacher-student continuation asymmetry on failed prefixes, where teachers redirect and students plow ahead, and convert it into a label-free handoff trigger, letting the teacher briefly generate a relay leg at detected trigger points before the student resumes and is optimized on the combined trajectory. With a Qwen3-4B-Instruct-2507 teacher and Qwen3-0.6B/1.7B students across eight math benchmarks, Relay-OPD places best or second-best on every benchmark, beating standard OPD by 5.73% and the strongest baseline FastOPD by 1.49% at 1.7B, while cutting training trajectory length by over 50%.
↳ Follow the thread