Fetching from the wire…
Public story · 2026-07-19 · high
263 task pairs across 42 sandboxes test whether an agent can tell a normal request from one it should refuse or escalate.
Why now: The benchmark is part of the AI research coverage dated July 19, 2026.
AgentAbstain scores AI agents on 263 task pairs spread across 42 executable sandbox environments, per the paper posted to arXiv. Each pair holds one task an agent should complete and a near-identical one it should refuse or escalate instead.
That second half is the one that costs money. Once an agent has write access to a filesystem, a database, or a payment API, the failure mode isn't finishing the wrong task. It's not knowing the task was wrong in the first place. An agent that completes 95% of tasks but never learns to say no is more dangerous than one that finishes 80% and knows its limits.
Nearly every agent benchmark I've seen scores the opposite problem: did the agent finish. AgentAbstain flips that by holding the task itself nearly constant across each pair, so the only thing being measured is judgment, not capability.
The paper doesn't say how frontier agents score on the refusal half, at least not in what's public in the arXiv listing. That's the number that would tell you whether this is a real, unsolved gap or something labs have quietly handled already.
Completion rate is the wrong leaderboard for anything with write access. If AgentAbstain catches on, watch for release notes that publish an abstention score next to the completion score, not just the completion number.
Each link below shares sources, entities, or timing with this story.
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (actually, agent, completion, cost, task); pushes against this story (versus).
Same source domain / Shared topic
Reported by the same outlet (arxiv.org); overlapping topics (actually, agent, benchmark, complete, completion).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (actually, agent, benchmark); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (agent, benchmark, task); pushes against this story (against).
Same source domain / Shared topic
Reported by the same outlet (arxiv.org); overlapping topics (agent, benchmark, cost, should).
Reported by the same outlet (arxiv.org); overlapping topics (access, agent, benchmark, task).
Reported by the same outlet (arxiv.org); overlapping topics (agent, benchmark, completion, task).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (agent, rate); pushes against this story (versus).