Agents
SkillEvo evolves agent skills from multi-turn interaction feedback, beating self-reflection by 23 points
arXiv 2608.13120 (2026-08-13) derives evolution gradients from multi-turn interaction feedback rather than from an agent's own post-hoc reflection, evaluated over six categories of cloud services, 9 production Skills and 98 skill-reference files. It surpasses self-reflection-based evolution by 23.0 points and single-turn-QA-driven evolution by 15.4 points. The claim underneath is that an agent grading its own transcript is a weak signal, and that the corrective information lives in what happened over subsequent turns — relevant to anyone maintaining a hand-tuned skill or SKILL.md library.
Source
↳ Follow the thread