Tool-calling leaderboards hide miscalibration because the state grader cannot see it
Aggregate accuracy on multi-turn tool-calling benchmarks averages over very different situations, and open-weight models now match or beat closed frontier models on that number. The authors decompose failures into action-class miscalibration and action-execution failure over a four-class space of TOOL_CALL, ASK, REFUSE and CONFIRM, and introduce a self-revealing bound Acc <= Gold Action Recall. Bound violation (Acc > GAR) exposes the state grader masking miscalibration, while large slack (GAR >> Acc) localizes execution failure inside TOOL_CALL. Across their model panel this separates heavily tool-trained families, whose standing the gap inflates, from families that pick a context-appropriate action.
Source
↳ Follow the thread