SimSkill Builds a Reusable Traffic-Simulation Skill Library and Gains 25 Points Without Touching the Backbone Weights
arXiv 2609.03753 builds a self-evolving agent around the SUMO traffic simulator that identifies its own capability gaps, generates and solves environment-grounded tasks, verifies solutions through an action-critic loop, and consolidates experience into episodic, procedural and semantic memory, all without updating the backbone model. Evaluated on two held-out benchmarks with three backbone LLMs under independent artifact-based verification, it improved verified completion by up to 25 percentage points, with ablations showing procedural and semantic memory contribute complementarily. The authors report the benefits are backbone- and budget-dependent and that memory does not improve every model or uniformly reduce inference cost, which is a useful caveat for anyone assuming a skill library is a free win.
↳ Follow the thread