Skills
Ranking your agent fleet by self-reported confidence can be worse than auditing at random — measure calibration against the δ* threshold first
For one human auditing N agents on a budget of B ≪ N per round, the paper derives a miscalibration threshold δ* past which confidence-ranked audit allocation underperforms uniform random selection, and shows δ* *rises* as the audit budget shrinks. Empirically, open-weight models produced near-constant, operationally useless confidence; only one proprietary model was calibrated enough to beat random. Actionable: before wiring confidence into a review queue, measure your models' calibration — if it exceeds δ*, randomize audits, because confidence-ranking is actively harmful, not merely neutral.
↳ Follow the thread