SEED: Self-Evolving On-Policy Distillation Pushes Agentic RL Training Without a Frozen Teacher
arXiv / HuggingFace Daily Papers·medium signal
arXiv 2607.14777 (104 upvotes on HuggingFace Daily Papers, July 17) proposes distilling agent policies on-policy from a teacher that itself evolves during training, rather than the standard fixed-teacher setup. For builders training task-specific agents, the pitch is that the teacher doesn't cap the student — a structural answer to why distilled agents typically plateau below their teacher on long-horizon tool use.