Skills
One training query recovers 71.5% of the states full-data on-policy distillation visits, and 16 queries match it outright
Training on-policy distillation with a single query keeps improving for hundreds of steps and recovers most of full-data OPD's gain across task domains and model families. Measuring state coverage, the fraction of full-data states a query set's rollouts reach, one query already hits 71.5% and 16 semantically distinct queries reach 98.9% and match full-data training. Alignment slows at the same rate either way, which the authors summarize as OPD being data-overfed and algorithm-starved. Content-light templates and off-domain WildChat queries also approached the real-query baseline.
↳ Follow the thread