Vibe Coding
RealSWE: 88% of real user prompts are bare problem statements, but only 7% of SWE-bench problems are
arXiv 2608.27831 (2026-08-28) builds a six-category information taxonomy and four style dimensions, then compares real prompts from SWE-chat against SWE-bench Verified and Pro. Requests carrying only a problem statement account for 88% of real prompts and 7% of benchmark problems, and 87% of real prompts are casually written versus 94% of benchmark problems written formally. The authors release 381 multi-variant tasks to close the gap, which matters for anyone reading SWE-bench numbers as a proxy for how an agent handles their own terse prompts.
Source
↳ Follow the thread