AgentAbstain Benchmarks the Skill Nobody Tests: Do LLM Agents Know When NOT to Act?
arXiv·medium signal
arXiv 2607.10059 introduces a paired-task benchmark of 263 task pairs across 42 executable sandbox environments, where each pair contains one task the agent should complete and a near-identical one it should refuse or escalate. Nearly every agent benchmark measures completion rate; this measures the inverse, which is the failure mode that actually matters when an agent has write access to a filesystem, a database, or a payment API.