Agents
SWE-Prometheus: coding agents asked to improve repo governance break existing behavior in up to 23% of runs
arXiv 2609.29465 gives agents a repository snapshot and an open-ended governance goal across 60 repos, then scores six dimensions with clean-environment probes and behavior gates. On the 22-repo public subset, ten models range from 0.0568 to 0.5760 mean Normalized Governance Improvement, with breakage rates from 0% to 23%. A template that ignores the repository scores 0.272 by adding tests, CI and docs, but improves reproducible environments and dependency security on none of the repos. The benchmark separates agents that add governance files from agents that make the build more reliable.
Source
↳ Follow the thread