Self-evolving agents hit a capability-contamination cliff: skill pools degrade past a critical size, and rollback doesn't fix it
Shang et al. (arXiv 2608.05810, submitted August 6) show that agents distilling reusable skills from their own trajectories improve only up to a critical pool size, after which new skills actively degrade performance — a defective skill becomes reference material for distilling later skills, forming cross-round contamination chains. The damage is structurally irreversible: deleting the source skill afterward recovers only a small fraction of the lost performance, because descendants already inherited the flawed reasoning. Their Verifier-as-Gatekeeper admits skills through three non-substitutable critics (structural validity, behavioral harmlessness, semantic consistency) plus marginal-gain subset selection, improving every round to 72% pass@1 on Terminal-Bench 2 with a pool roughly 5x smaller, and the frozen pool transfers to four other backbones without re-evolution.
Source
↳ Follow the thread