SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions
arXiv·high signal
A new testbed that evaluates coding agents on multi-turn, interactive, user-driven software-engineering sessions instead of single-shot patch generation. This is closer to how developers actually use coding agents day-to-day, exposing failure modes that static SWE-bench-style evals miss.