Stop your research agent on evidence sufficiency, not search count: answer accuracy tracks cumulative retrieval recall, not effort
Using human-annotated document-level relevance judgments to score the evidence retrieved at each step, this trajectory-level diagnosis of six deep-search agents on BrowseComp-Plus and BrowseComp separates retrieval gaps (the evidence was never found) from utilization gaps (it was found and misused). Search effort and answer quality are only weakly aligned: accuracy correlates with cumulative retrieval recall far more than with number of searches or context consumed, useful evidence usually appears early in the trajectory, and agents keep searching anyway — producing a long tail of low-yield steps. Exploratory query reformulation stays useful, but the best-performing agents issue far fewer redundant queries, which argues for an explicit stopping criterion based on whether sufficient supporting evidence has been gathered rather than a fixed step or token budget.
↳ Follow the thread