Research
AppWorld-UL: Claude Opus 4.7 Scores 48.6% When the Agent Has to Talk to the User, 21.3% at Scenario Level
Chen, Trivedi, Pan, Zhang, Srinivasan, Balasubramanian, and Sabharwal extend AppWorld into a 516-task 'user-in-the-loop' benchmark across nine simulated apps (Amazon, Spotify, and others), deliberately injecting ambiguities and constraints that force the agent to ask clarifying questions, request confirmation, or declare a task infeasible. Claude Opus 4.7 manages only 48.6% success overall, 35.7% on the compositional subset, and 21.3% under the stricter scenario-level metric. User behavior is simulated by an LLM with designed knowledge boundaries rather than the unconstrained simulators prior work used — and the analysis finds correct interaction, not tool execution, is what determines success.
↳ Follow the thread