Reddit
Study of 4,882 Agent-Authored PRs: Half Ship No Tests, and Error Handling Goes Untested 86% of the Time
Dipongkor, Baral, Lam and Moran (arXiv:2607.18057, July 20, accepted to ICSME 2026) analyzed 4,882 pull requests from five coding agents in the AIDev dataset — 532 Java, 4,350 Python. Agents modify tests in only 49.6% of PRs that touch testable code; existing suites cover just 61.5% of changed lines in Java and 27.0% in Python; and when agents do write tests, coverage actually improves in only 35.9% (Java) and 22.5% (Python) of submissions. Error-handling constructs are the worst blind spot, with miss rates of 86.0% in Java and 81.0% in Python. The practical takeaway: reviewing agent PRs on diff quality alone systematically misses the untested failure paths.
↳ Follow the thread