Asking a Model How Confident It Is Gives Rank Information But Not Calibration, Across 30 Models and 10 Tasks
arXiv 2608.28382 (2026-08-28, cs.AI/cs.CL) compares verbalized confidence against logits-based confidence on 8 classification tasks and semantic-entropy uncertainty on 2 generation tasks, spanning 30 models from three families. Instance-level association is weak on average and improves only on easier items and stronger base models; instruction-tuned models report higher confidence and sometimes higher association but show larger confidence gaps and worse calibration. Prompt design mostly shifts the reported distribution rather than the alignment, with attitude cues inflating confidence without improving it, supporting a lossy-channel view where linguistic confidence needs multi-axis diagnostics before it feeds any downstream reliability pipeline.
↳ Follow the thread