Research
Coding Agents Solve Only 51.2% of Real Dependency Upgrades That Hide Breaking Changes
DEPBENCH collects 203 real-world dependency-upgrade tasks across five package ecosystems and five language communities, each containing code-level breakage that the upgrade never communicated to maintainers and that requires source adaptation. The best completed agent configuration solves 104 of 203 tasks, 51.2%, with wide variation across agent harnesses, models and ecosystems. This is the maintenance half of software work that SWE-bench-style benchmarks mostly skip, and it puts a number on the gap between agent capability and everyday upgrade toil.
↳ Follow the thread