Fetching from the wire…
Public story · 2026-08-31 · high
A new paper finds 88% of actual developer requests are bare problem statements, versus 7% of the benchmark problems models get graded on.
Why now: The paper's 381 new tasks, posted to arXiv in August 2026, are built to close that exact gap.
A paper posted to arXiv in August 2026 argues the industry has been grading coding agents on the wrong kind of question. RealSWE built a six-category taxonomy of what information a bug report contains, plus four style dimensions, then ran it against real prompts pulled from SWE-chat and compared them to the problems in SWE-bench Verified and Pro.
The gap is stark. 88% of real prompts are bare problem statements: no repro steps, no stack trace, no suggested fix, just "this is broken." Only 7% of SWE-bench problems look like that. On the style side, 87% of real prompts read as casual, dashed off the way a developer actually talks. 94% of the benchmark prompts are formally written, closer to a ticket a technical writer polished than a note a tired engineer fired off at 11pm.
That mismatch matters because SWE-bench scores are the number everyone quotes when comparing coding agents. If the benchmark prompt is unusually complete and unusually formal, a model can look sharp at parsing well-specified tickets while never being tested on the ambiguous, underspecified request that shows up in an actual issue tracker. The paper's answer is 381 multi-variant tasks built to cover the real distribution instead of the sanitized one.
What the paper doesn't say is how much of the SWE-bench score gap is style versus real information loss. A bare problem statement might just take an extra clarifying turn, or it might strip out details a model needs and can't recover on its own. RealSWE gives the taxonomy to measure that. It doesn't yet say which failure mode dominates.
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Both cover SWE, Verified; reported by the same outlet (arxiv.org); overlapping topics (against, compar, problem).
Shared entities / Same source domain / Shared topic / Earlier coverage
Both cover SWE, Verified; reported by the same outlet (arxiv.org); overlapping topics (benchmark, swe bench).
Shared entities / Same source domain / Earlier coverage / Tension
Both cover SWE, Verified; reported by the same outlet (arxiv.org); earlier SWE coverage from 2026-08-24.
Shared entities / Same source domain / Earlier coverage / Downstream implication
Both cover SWE, Verified; reported by the same outlet (arxiv.org); earlier SWE coverage from 2026-08-19.
Shared entities / Same source domain / Earlier coverage / Tension
Both cover SWE, Verified; reported by the same outlet (arxiv.org); earlier SWE coverage from 2026-07-23.
Shared entities / Shared topic / Earlier coverage / Tension
Both cover SWE, Verified; overlapping topics (against, benchmark); earlier SWE coverage from 2026-07-17.
Shared entities / Same source domain / Earlier coverage
Both cover SWE, Verified; reported by the same outlet (arxiv.org); earlier SWE coverage from 2026-08-30.
Both cover SWE, Verified; reported by the same outlet (arxiv.org); earlier SWE coverage from 2026-08-28.