Isotonic Calibration of Agent Confidence Raises Committed-Answer Accuracy 41 Points but Cuts Overall Accuracy 17 Points on MuSiQue
arXiv 2608.26846 proposes matched trajectory replay, holding candidate answer states, evidence points, budgets and action costs fixed so confidence-to-action mappings can be compared by their trajectory-level consequences rather than in isolation. Across Mistral, GPT and Qwen on HotpotQA and MuSiQue, post-hoc isotonic calibration at the same numerical threshold raises accuracy among committed answers by up to 41 points in all six pairs, but overall accuracy improves up to 15 points on HotpotQA while falling up to 17 points on MuSiQue, and a calibration map fitted before retrieval is worse than raw confidence at retrieval depth three. The conclusion is that calibration makes commitment risk interpretable but does not estimate the value of another retrieval, which needs a separate utility estimate.
↳ Follow the thread