arXiv: Internalizing Agency from Reflective Experience — Self-Correcting LLM Agents Without External Supervision
arXiv·medium signal
A new paper proposes training LLMs to internalize agency through reflective experience, enabling agents to plan, act, and self-correct from multi-step mistakes during long-horizon interactions without requiring human-provided supervision signals. The approach accumulates in-context failure patterns and uses them to update agent behavior within the same session. Results show improved recovery from compounding errors compared to standard agent architectures that treat each step as independent.