Fetching from the wire…
Research2026-08-06 · source-backed
arXiv 2608.04682 removes the assumption that every SWE benchmark makes, that a high-quality issue report exists. Six bug categories, eight languages, multi-bug fixing and potential-bug discovery under dual-track evaluation. Most state-of-the-art coding agents perform poorly at locating recorded bugs without report guidance, handling multi-bug scenarios, and surfacing valid potential bugs. That's the gap between SWE-bench-style scores and what happens when you point an agent at a repo with no ticket, which is most of real work.
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source domain / Shared topic / Earlier coverage
Both cover Most, SWE; reported by the same outlet (arxiv.org); overlapping topics (agent, benchmark, between).
Shared entities / Shared topic / Earlier coverage
Both cover Most, SWE; overlapping topics (agent, benchmark, coding, evaluation); earlier Most coverage from 2026-04-12.
Shared entity: SWE / Same source domain / Shared topic / Earlier coverage / Tension
Both cover SWE; reported by the same outlet (arxiv.org); overlapping topics (agent, benchmark, evaluation).
Shared entity: SWE / Same source domain / Shared topic / Earlier coverage
Both cover SWE; reported by the same outlet (arxiv.org); overlapping topics (agent, benchmark, coding, issue).
Shared entity: SWE / Same source domain / Shared topic / Earlier coverage / Tension
Both cover SWE; reported by the same outlet (arxiv.org); overlapping topics (agent, coding).
Shared entity: SWE / Shared topic / Earlier coverage / Tension
Both cover SWE; overlapping topics (agent, benchmark, category, coding); earlier SWE coverage from 2026-02-12.
Shared entity: SWE / Same source domain / Shared topic / Earlier coverage
Both cover SWE; reported by the same outlet (arxiv.org); overlapping topics (agent, benchmark, coding).
Both cover SWE; reported by the same outlet (arxiv.org); overlapping topics (agent, benchmark, coding).