ContinualSkillBench finds explicit skill libraries roughly tie plain in-context learning — the gains come from context, not skills
Testing in-context continual skill learning across five domains of 100 interconnected subtasks each, ordered by rising difficulty with deliberate cross-task reuse opportunities, the benchmark found sequential execution helps but that in-context learning performs comparably to explicit skill maintenance on average. That suggests most improvement comes from adaptation to prior context and feedback rather than genuine reusable skill abstraction; explicit skills only pay off selectively, for tasks needing reusable procedures or precise outputs. Weaker models accumulated larger, more fragmented collections of task-specific skills — an argument for auditing skill-library growth rather than treating it as progress.
Source
↳ Follow the thread