Coding Agents Burn 631K Tokens Per Resolved SWE-Bench Issue Just Finding the File — CodeGrep Shows Retrieval Only Helps Above 0.677 Precision
CodeGrep (arXiv 2608.05886, Aug 6) measures that a 30B OpenHands agent averages 23 rounds and 631K tokens per resolved SWE-Bench Verified issue, much of it spent on grep, glob, and view_file during exploration. A 14B retrieval agent trained end-to-end with GRPO raises resolve rate to 27.0% from 25.8% while cutting 15% of rounds and 19% of tokens on resolved instances. The sharpest finding is a precision threshold for downstream utility: BM25 at 0.375 precision actively degrades the agent, Jina at 0.445 is neutral, and only CodeGrep at 0.677 helps — meaning bolting a mediocre retriever onto your coding agent makes it worse, not slower-but-better. Model, training pipeline, RL environment, and eval harnesses are promised for release.
↳ Follow the thread