Research
Lightning OPD 2.0 Tackles Style Bias When the Distillation Teacher Didn't Write the SFT Data
On-policy distillation gives dense token-level supervision but its effectiveness depends on teacher consistency — the model providing OPD supervision should also have generated the SFT demonstrations. That condition is routinely violated in practice when SFT data has mixed or unknown provenance, or when teams deliberately pick one model for data generation and a different one for distillation. Lightning OPD 2.0 targets the resulting cross-teacher style bias in large reasoning models. Directly relevant to anyone assembling reasoning-model training pipelines from heterogeneous open datasets, which is most of them.
Source
↳ Follow the thread