Fetching from the wire…
Public story · 2026-08-25 · high
Run each task twice and the pass rate for retail, insurance and IT support workflows drops to 25%.
Why now: The benchmark paper posted in the August 25 research roundup.
A new benchmark called Thinkingbox scores AI agents on 507 business workflows, spanning retail, hospitality, auto insurance, neobank internal IT and consulting support, and checks the results against the actual state of a backend system rather than just the agent's own report of success, according to the paper. The best agent tested hit 65.36% pass@1, meaning it completed the task correctly on a single try about two-thirds of the time.
Run the same task twice and require both attempts to succeed, and that number falls to 25.25%. That's the gap between an agent looking competent in a demo and an agent you'd trust to run unattended.
The design is what makes the drop meaningful instead of just pessimistic. Thinkingbox runs each task in an isolated sandbox with full execution traces, then checks outcomes against terminal backend state, an approach the paper says accepts valid paths to the goal while rejecting trajectories that leave wrong, missing, or extra effects behind. An agent that books the right refund through an unexpected sequence of steps still passes. One that leaves a stray charge on the account doesn't, even if it reports success.
That distinction matters more than the headline number for anyone deciding what to automate. A 65% single-shot score sounds usable for something with a human checking the output. A 25% two-run consistency score is a different conversation, especially in domains like insurance claims or bank IT tickets where an unnoticed wrong effect isn't a UI glitch, it's a real transaction. The paper doesn't say which specific failure modes account for the gap between the two numbers, so it's not yet clear whether agents are failing randomly or failing in ways tied to specific task types.
Each link below shares sources, entities, or timing with this story.
Claude Code uses MCP / Shared entity: MCP / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses MCP); both cover MCP; reported by the same outlet (arxiv.org).
Claude Code uses MCP / Shared entity: MCP / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Code uses MCP); both cover MCP; reported by the same outlet (arxiv.org).
OpenCode supports MCP / Shared entity: MCP / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (OpenCode supports MCP); both cover MCP; overlapping topics (against, agent, around).
Claude Code uses MCP / Shared entity: MCP / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses MCP); both cover MCP; overlapping topics (against, agent).
Claude Code uses MCP / Shared entity: Best / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses MCP); both cover Best; reported by the same outlet (arxiv.org).
Claude Code uses MCP / Shared entity: MCP / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses MCP); both cover MCP; overlapping topics (against, agent).
Claude uses MCP / Shared entity: MCP / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude uses MCP); both cover MCP; overlapping topics (against, agent).
Claude Code uses MCP / Shared entity: MCP / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses MCP); both cover MCP; overlapping topics (against, agent).