Fetching from the wire…
Public story · 2026-07-31 · high
Only one of the tested models was calibrated enough for confidence-ranking to beat random picks, and the failure threshold climbs as budgets shrink.
Why now: The paper's calibration findings surfaced in coverage on July 31.
A new arXiv paper derives the point where sorting agent audits by self-reported confidence turns worse than picking cases at random, per the study.
That matters for anyone routing agent output through a review queue. Miscalibrated confidence scores can send reviewers to the wrong cases more often than a coin flip would.
The threshold isn't fixed, either. It climbs higher as your audit budget shrinks, per the paper.
Empirically, that cushion didn't help much. Open-weight models tested produced confidence scores that stayed nearly constant, regardless of whether the output was right or wrong. Only one proprietary model tested was calibrated enough for ranking to beat random selection at all.
The fix is simple, per the paper: measure calibration before letting a model's confidence score decide what a human reviews first. Past that threshold, the recommendation is blunt. Randomize the queue, because ranking there isn't neutral. It's actively worse than guessing.
Each link below shares sources, entities, or timing with this story.
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (agent, beat, budget); pushes against this story (but).
Same source domain / Shared topic
Reported by the same outlet (arxiv.org); overlapping topics (agent, allocation, beat, budget).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (beat, model); pushes against this story (versus).
Reported by the same outlet (arxiv.org); overlapping topics (agent, model); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (agent, model); pushes against this story (but).
Same source domain / Shared topic / Downstream implication
Reported by the same outlet (arxiv.org); overlapping topics (audit, model); traces where this leads (downstream).
Reported by the same outlet (arxiv.org); overlapping topics (agent, model); traces where this leads (downstream).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (agent, beat); pushes against this story (but).