An evolving skill library lifts a fixed computer-use stack 5.7 to 18.6 points, but repeated edits do not always recover the original task
This online skill-evolution framework turns interaction trajectories and evaluator feedback into a persistent versioned library, with each iteration executing against a frozen snapshot so evidence-guided updates only land in later iterations and no model parameters change. Compared against a configuration-matched empty-library control across four OSWorld application domains with identical action-generation and grounding stacks, the evolving library won all four post-warm-up domain runs by 5.7 to 18.6 percentage points. Provenance analysis in GIMP found skills retrieved across task-of-origin boundaries and revision churn where repeated accepted edits failed to recover the originating task, so the gains are conditional rather than monotone. Code is at github.com/LongtaoHu/Skill-Evo4GUI.
Source
↳ Follow the thread