Fetching from the wire…
Research2026-07-28 · source-backed
Models don't treat 0-100 as continuous. Verbalized confidence is badly overconfident (Xiong et al.) and only calibrates after task-specific fine-tuning (Lin et al.). The recursion problem is the core of it: if you can't trust the answer, you can't trust the model's confidence in its confidence. The recommended alternative is semantic entropy (Farquhar et al.). Note this post is a synthesis, not original experiments. (justinflick.com)
Each link below shares sources, entities, or timing with this story.
Simon Willison released LLM / Shared entity: LLM / Shared topic / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; overlapping topics (answer, model).
Simon Willison released LLM / Shared entity: Models / Shared topic / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover Models; overlapping topics (model, only).
LLM uses OpenAI / Shared entity: LLM / Earlier coverage / Tension
Linked by a graph relationship (LLM uses OpenAI); both cover LLM; earlier LLM coverage from 2026-07-27.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-18.
LLM uses OpenAI / Shared entity: LLM / Earlier coverage
Linked by a graph relationship (LLM uses OpenAI); both cover LLM; earlier LLM coverage from 2026-06-19.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-07-19.
Simon Willison released LLM / Shared entity: Note / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover Note; earlier Note coverage from 2026-07-16.