Feed your self-evolving agent its failures, not its wins: all 11 skill selections that beat baseline came from feedback conditions containing failed trajectories
A controlled study of 42 feedback runs across 14 model-benchmark settings varied whether evolving skills were shown successes, failures, or both. Evolution turns out to be sparse — only 55 of 388 candidates produced a byte-distinct validation best — and validation-based selection picked an evolved skill in 11 of 14 settings, 9 of which improved test performance. Every single one of those 11 selections came from a condition that included failed trajectories, which inverts the common practice of curating a success-only skill library; the authors also caution that returns are model- and benchmark-dependent (GPT-5.5 oracle sampling closed to 0.43 points on SearchQA but stayed 30.96 points behind on SpreadsheetBench), so this is search under validation, not steady improvement from more rounds.
↳ Follow the thread