SDAR: Self-Distilled Agentic Reinforcement Learning — Token-Level Gating Yields +9.4% on ALFWorld, +10.2% on WebShop Over GRPO
arXiv·medium signal
ZJU researchers publish SDAR (arXiv 2605.15155, May 14): treats on-policy self-distillation as a gated auxiliary objective while keeping RL as primary backbone. A sigmoid gate lets each token decide its own supervision intensity — strengthening distillation on teacher-endorsed tokens, attenuating rejections. Consistently outperforms GRPO and hybrid baselines across Qwen2.5/Qwen3 families without instability.