Active-SWE: 1,663 Tasks Asking Coding Agents to Find Bugs With No Issue Report, and Most SOTA Agents Struggle
Existing SWE benchmarks assume a high-quality issue report always exists, which rarely holds in practice; Active-SWE removes it, covering 1,663 tasks across six bug categories and eight languages where agents must proactively discover and fix multiple bugs without report guidance. It uses a difficulty-aware task formulation pipeline and dual-track evaluation, expanding scope from fixing one recorded bug to multi-bug fixing and potential-bug discovery. The reported result is that most state-of-the-art coding agents perform poorly on locating and resolving recorded bugs, handling multi-bug scenarios, and surfacing valid potential bugs — a gap between SWE-bench-style scores and how agents behave when pointed at a repo with no ticket.
↳ Follow the thread