Fetching from the wire…
Public story · 2026-08-25 · high
The paper's fix, NIS-Agent, isolates context at two points and cuts token cost with no drop in accuracy.
Why now: As of August 25, the paper is the newest evidence that agent context design directly affects both accuracy and token cost.
A new benchmark called IBIS holds an agent's search results fixed and varies only who wrote the step before them, then scores worse when the agent did. Researchers call the effect inertia bias: an agent grades its own plans more leniently, even when the evidence hasn't changed.
The bias shows up in how an agent judges the consequences of a query or plan it already produced. IBIS isolates that judgment by holding the search observations constant, so the accuracy gap traces only to authorship of that earlier step.
NIS-Agent, described in the same paper, walls off the context that produced a step from the context that judges it. It does this at two points: triaging which webpages to trust, and validating the agent's own final answer. Splitting judgment from authorship at just those two spots holds accuracy on GAIA, WebWalkerQA and BrowseComp. Token use drops by a third.
Each link below shares sources, entities, or timing with this story.
Anthropic benchmarked against GAIA / Same source domain / Shared topic
Linked by a graph relationship (Anthropic benchmarked against GAIA); reported by the same outlet (arxiv.org); overlapping topics (agent, context).
Anthropic benchmarked against GAIA / Shared entity: GAIA / Shared topic / Earlier coverage
Linked by a graph relationship (Anthropic benchmarked against GAIA); both cover GAIA; overlapping topics (agent, context, model, token).
Anthropic benchmarked against GAIA / Shared entity: Agent / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Anthropic benchmarked against GAIA); both cover Agent; overlapping topics (agent, cost).
Anthropic benchmarked against GAIA / Shared entity: BrowseComp / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Anthropic benchmarked against GAIA); both cover BrowseComp; overlapping topics (cost, model).
Shared entity: Agent / Same source domain / Shared topic / Earlier coverage
Both cover Agent; reported by the same outlet (arxiv.org); overlapping topics (action, agent, benchmark, cost, model).
Shared entities / Same source domain / Shared topic / Earlier coverage
Both cover BrowseComp, GAIA; reported by the same outlet (arxiv.org); overlapping topics (agent, model).
Anthropic benchmarked against GAIA / Shared entity: Agent / Shared topic / Earlier coverage
Linked by a graph relationship (Anthropic benchmarked against GAIA); both cover Agent; overlapping topics (action, agent, model).
Anthropic benchmarked against GAIA / Shared entity: BrowseComp / Shared topic / Earlier coverage
Linked by a graph relationship (Anthropic benchmarked against GAIA); both cover BrowseComp; overlapping topics (browsecomp, context, cost).