Fetching from the wire…
Public story · 2026-07-26 · high
AppWorld-UL gives the simulated user real knowledge gaps, and the paper traces the model's failures to conversation, not tool calls.
Why now: The benchmark's numbers come from arXiv paper 2607.20536, current as of July 26, 2026.
Claude Opus 4.7 succeeds on just 21.3% of tasks in AppWorld-UL, a benchmark that forces clarifying questions, per the researchers who built it.
That's a benchmark built for agents that have to work with users who don't already know what they want. Most agents deployed in production never get tested against that kind of ambiguity.
AppWorld-UL extends AppWorld into 516 tasks across nine simulated apps, including Amazon and Spotify, authored by Chen, Trivedi, Pan, Zhang, Srinivasan, Balasubramanian, and Sabharwal. Every task is seeded with ambiguity on purpose. The agent has to ask a clarifying question, request confirmation, or say a task can't be done, instead of calling tools until something works.
Overall, Opus 4.7 succeeds on 48.6% of tasks. Narrow to the compositional subset, tasks that chain multiple actions together, and it drops to 35.7%. The 21.3% figure is the stricter scenario-level score, where every step in a multi-turn interaction has to land correctly.
Earlier work let an LLM stand in for the user, answering nearly anything the agent asked. That let an agent route around its own confusion just by asking better questions. AppWorld-UL's simulated user has designed knowledge boundaries instead. Sometimes the honest answer is 'I don't know,' and the agent has to act on that rather than extract information that isn't there.
The paper's analysis traces failure to interaction, not tool execution. Agents lose points recognizing an ambiguous request, asking the right follow-up, or working with a user who can't fully specify what they want.
A related benchmark tests coding agents on a different axis. It's built from real commits and rewritten so the prompt can't be web-searched, checking whether an agent can solve a problem without memorizing the answer. AppWorld-UL asks a narrower question: can the agent tell when it doesn't have enough to go on, and say so instead of guessing.
Each link below shares sources, entities, or timing with this story.
Sandbox memory in the test suite peaks at 28 GB a session, and latency across components swings up to 32x within the same app.
ActBench ran 24,000 attack trajectories across 15 models and six harnesses; no harness pushed success below 73.7%.
Chen et al. extend AppWorld into a 516-task user-in-the-loop benchmark across nine simulated apps, injecting ambiguities and constraints that force the agent to ask clarifying questions, request confirmation, or declare a task infeasible (arXiv 2607.20536). Opus 4.7 gets 48.6%...
The technique filters terminal output an agent reads one line of, and resolution rates held steady on 50 SWE-bench Lite tasks.
Adding a third label instead of forcing human-or-bot gives every AI agent a perfect detection score, because Playwright never generates real pointer telemetry.
FrontierChallenge tested 12 models on 97 lab workflows, and Claude Code claimed success in 75.5% of the runs it failed.
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.