Internalizing Agency from Reflective Experience: Fixing Distribution Sharpening in RL-Trained Agents
arXiv 2603.16843·high signal
Standard RL post-training with verifiable rewards causes 'distribution sharpening' — agents get better at reproducing a narrow set of already-successful trajectories while ignoring rich environment feedback signals. This paper proposes training agents to internalize agency from reflective experience, using all environment feedback (not just final success signals) to build more generalizable planning and error-recovery behaviors. The method targets long-horizon autonomous agents submitted to ICML 2026.