SWE-Touch Shows Coding Agents Lose 7.7 Points of Resolve Rate the Moment a Human Edits the Same Files
SWE-Touch stress-tests the shared-workspace case that repository benchmarks ignore: it mines task-critical regions from repair trajectories, uses a separate User Patch Generator to build plausible 'Counter-Edits' that conflict with task completion, and injects them with contextual user messages when the agent reaches that code. Across nine coding models, Counter-Edits drop average resolve rate on SWE-bench Verified by 7.7 percentage points, with degradation persisting on SWE-Bench Pro and DeepSWE. Trajectory analysis traces the failures to agents retaining conflicting code or overwriting it without re-inspecting the repo or running targeted tests — a direct warning for anyone editing files while an agent is mid-task.
↳ Follow the thread