Research
LLMs Encode Their Own Knowledge Boundary and Still Choose the Specific Wrong Answer Over a Correct Vague One
arXiv 2608.13484 (2026-08-13) frames hallucination through Grice: an uncertain cooperative speaker retreats up the specificity hierarchy, trading informativeness for truthfulness. Using a T-REx-based benchmark varying entity familiarity and referent specificity, the authors show model activations do encode whether a referent falls inside the knowledge boundary, and models do anticipate the specificity of what they are about to generate — but the two signals are never reconciled at generation time. Models overwhelmingly prefer specific referents for unknown entities even when correct generic alternatives are offered on a plate. The substrate for abstention exists; the policy does not.
↳ Follow the thread