Fetching from the wire…
Public story · 2026-08-31 · high
Testing 30 models found the confidence a model states out loud barely lines up with whether its answer is right, and fine-tuning widens the gap.
Why now: The paper posted August 31, 2026.
Researchers tested whether a model's own stated confidence, the number you get when you ask how sure it is, actually predicts whether it's right. Across 30 models from three families, they compared that self-report against two harder benchmarks: logits-based confidence on 8 classification tasks, and semantic entropy on 2 generation tasks.
The self-reported number ranks answers roughly right. More often than not, correct answers get higher confidence than wrong ones. But that ranking is weak on average, and it only firms up on easier questions and with stronger base models. On harder tasks or weaker models, the self-report and actual correctness barely track each other.
Instruction-tuned models make this worse. They report higher confidence overall, and sometimes rank answers slightly better, but the gap between what they claim and what they deliver widens. Calibration gets worse even as the ranking gets marginally better. That cuts against the assumption that tuning makes a model more trustworthy about its own limits.
Prompt design doesn't fix it either. Changing how you ask for confidence shifts the distribution of numbers a model reports, not how well those numbers line up with reality. Cueing the model with attitude, framing the question as if confidence matters, inflates the reported number without making it more accurate.
The paper's own recommendation is narrow. Use self-reported confidence to sort candidate answers against each other, never as a cutoff for auto-approving or rejecting one. An app with a slider or badge showing a model's confidence, letting users treat 90 percent as a green light, is running on a mismatch this paper measured directly. The gap doesn't close by tuning the model further. It's built into what self-report is.
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source domain / Shared topic / Tension
Both cover Instruction, Prompt; reported by the same outlet (arxiv.org); overlapping topics (average, model).
Shared entity: Prompt / Same source domain / Shared topic / Earlier coverage / Tension
Both cover Prompt; reported by the same outlet (arxiv.org); overlapping topics (model, task).
Both cover Prompt; reported by the same outlet (arxiv.org); overlapping topics (classification, model).
Shared entity: Asking / Same source domain / Shared topic / Earlier coverage
Both cover Asking; reported by the same outlet (arxiv.org); overlapping topics (asking, model).
Shared entity: Prompt / Shared topic / Earlier coverage / Tension
Both cover Prompt; overlapping topics (against, model); earlier Prompt coverage from 2026-06-28.
Shared entity: Asking / Shared topic / Earlier coverage
Both cover Asking; overlapping topics (asking, confidence, model); earlier Asking coverage from 2026-07-20.
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (against, average, model); pushes against this story (against).
Same source domain / Shared topic
Reported by the same outlet (arxiv.org); overlapping topics (confidence, confident, model, task).