Fetching from the wire…
Public story · 2026-08-16 · high
The paper's own benchmark shows models choosing a specific wrong answer over a correct generic one already sitting in reach.
Why now: The paper posted to arXiv in August 2026.
Large language models encode exactly when they're uncertain about an answer, then pick a specific wrong guess anyway, per a paper posted to arXiv.
That gap matters for anyone building on model output: hallucination on unfamiliar entities isn't a missing-signal problem. The model already has the right signal and doesn't act on it.
The researchers built a benchmark on T-REx. It varies how familiar an entity is to the model and how specific the correct answer needs to be.
Model activations encode both signals separately, per the paper. One tracks whether a fact sits inside the model's knowledge boundary. The other tracks how specific the response it's about to generate will be.
Those two signals never talk to each other at generation time. Faced with an unfamiliar entity, models overwhelmingly reach for a specific, wrong referent. A correct generic answer sits in the same output space and goes unused.
The paper frames this in Gricean terms. A cooperative speaker who's uncertain trades informativeness for truthfulness, retreating up a specificity hierarchy instead of guessing wrong and specific.
These models don't make that trade. The substrate for abstention exists in the activations, but the policy doesn't.
Each link below shares sources, entities, or timing with this story.
Shared entity: LLMs / Same source domain / Shared topic / Earlier coverage / Tension
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (benchmark, model).
Shared entities / Shared topic / Earlier coverage
Both cover LLMs, Models; overlapping topics (knowledge, model); earlier LLMs coverage from 2026-03-20.
Shared entity: Models / Same source domain / Shared topic
Both cover Models; reported by the same outlet (arxiv.org); overlapping topics (abstention, answer, benchmark).
Shared entity: Models / Same source domain / Shared topic / Earlier coverage
Both cover Models; reported by the same outlet (arxiv.org); overlapping topics (cooperative, model).
Shared entities / Earlier coverage / Tension
Both cover LLMs, Models; earlier LLMs coverage from 2026-07-31; pushes against this story (against).
Shared entity: Models / Same source domain / Shared topic / Earlier coverage
Both cover Models; reported by the same outlet (arxiv.org); overlapping topics (benchmark, model).
Shared entity: LLMs / Same source domain / Shared topic / Earlier coverage
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (benchmark, model).
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (answer, model).