Fetching from the wire…
Public story · 2026-07-26 · high
AppWorld-UL gives the simulated user real knowledge gaps, and the paper traces the model's failures to conversation, not tool calls.
Why now: The benchmark's numbers come from arXiv paper 2607.20536, current as of July 26, 2026.
Claude Opus 4.7 succeeds on just 21.3% of tasks in AppWorld-UL, a benchmark that forces clarifying questions, per the researchers who built it.
That's a benchmark built for agents that have to work with users who don't already know what they want. Most agents deployed in production never get tested against that kind of ambiguity.
AppWorld-UL extends AppWorld into 516 tasks across nine simulated apps, including Amazon and Spotify, authored by Chen, Trivedi, Pan, Zhang, Srinivasan, Balasubramanian, and Sabharwal. Every task is seeded with ambiguity on purpose. The agent has to ask a clarifying question, request confirmation, or say a task can't be done, instead of calling tools until something works.
Overall, Opus 4.7 succeeds on 48.6% of tasks. Narrow to the compositional subset, tasks that chain multiple actions together, and it drops to 35.7%. The 21.3% figure is the stricter scenario-level score, where every step in a multi-turn interaction has to land correctly.
Earlier work let an LLM stand in for the user, answering nearly anything the agent asked. That let an agent route around its own confusion just by asking better questions. AppWorld-UL's simulated user has designed knowledge boundaries instead. Sometimes the honest answer is 'I don't know,' and the agent has to act on that rather than extract information that isn't there.
The paper's analysis traces failure to interaction, not tool execution. Agents lose points recognizing an ambiguous request, asking the right follow-up, or working with a user who can't fully specify what they want.
A related benchmark tests coding agents on a different axis. It's built from real commits and rewritten so the prompt can't be web-searched, checking whether an agent can solve a problem without memorizing the answer. AppWorld-UL asks a narrower question: can the agent tell when it doesn't have enough to go on, and say so instead of guessing.
Each link below shares sources, entities, or timing with this story.
Same source
Cite the same source (arXiv 2607.20536).
Same source domain / Semantically similar
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.73).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.70).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.68).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.68).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.67).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.66).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.66).