1% of tokens can be enough for on-policy distillation when selection accounts for gradient noise
arXiv / HuggingFace Daily Papers·low signal
arXiv 2609.24432 (21 Sep) proposes an information-efficiency ratio (IER) that scores tokens by both teacher usefulness and how reliably their gradient can be estimated. On math and medical reasoning, sparse distillation at 0.1% to 1% token budgets matches or beats full on-policy distillation without token selection. Code is at github.com/BruceSheng1202/IER-OPD. Single paper, not yet replicated.