Skills
Distilling 1,000 ML repos into 5,000 verified skills lifted an agent 134% on MLE-bench with the model and budget held fixed
Repo-To-Skill (arXiv 2609.02749, submitted 2026-09-02) argues the missing layer in research agents is operational know-how, and extracts it from existing GitHub repositories into the AREX-Skill Library: 5,000+ verified skills from 1,000 widely used ML repos, organized into 20 areas and 178 capability families. With the GPT-5.5 backbone, harness, and execution budget held constant, the skill-equipped agent scored 134.3% higher on MLE-bench, 34.4% higher on PaperBench, 9.2% higher on FrontierCS, and 14.0% higher on PassNet. For builders this is the clearest evidence yet that mining your own dependency repos into skill files is a cheaper capability lever than swapping models.
↳ Follow the thread