Fetching from the wire…
Public story · 2026-08-16 · high
A hidden-state probe predicts answerability with 0.91 AUROC, but no model scores above 0.292 on the benchmark.
Why now: TRAPSBench posted to arXiv in August 2026, and the paper doesn't explain why models catch textual impossibility about 4 times more easily than missing visual evidence.
TRAPSBench pits 16 vision-language models against 1,404 matched physics video pairs, each one rigged so a single change makes the answer undeterminable. The best model scored 0.292 on a metric built to punish wrong guesses as hard as it rewards right ones, per the arXiv paper. These systems answer confidently even when the video hides the evidence they'd need.
A linear probe run against the models' hidden states predicted whether a question was even answerable, up to 0.91 AUROC, the researchers found. The model's internals knew. Its output didn't say so.
Steering makes the case stronger. Nudging a single "void" direction in one layer of the network causally turns abstention on or off, the paper reports.
Models catch a textually impossible question about 4 times more often than they catch a video missing the same visual evidence. Language gives them an out. Pixels don't.
That looks less like confusion than reluctance. A model that decodes answerability at 0.91 AUROC but scores only 0.292 already has the signal. It just doesn't act on it before answering.
Each link below shares sources, entities, or timing with this story.
Shared entity: Models / Same source domain / Shared topic
Both cover Models; reported by the same outlet (arxiv.org); overlapping topics (abstention, answer, benchmark).
Shared entity: Models / Same source domain / Shared topic / Earlier coverage
Both cover Models; reported by the same outlet (arxiv.org); overlapping topics (benchmark, best).
Shared entity: AUROC / Same source domain / Earlier coverage / Tension
Both cover AUROC; reported by the same outlet (arxiv.org); earlier AUROC coverage from 2026-08-14.
Both cover AUROC; reported by the same outlet (arxiv.org); earlier AUROC coverage from 2026-08-04.
Shared entity: Models / Same source domain / Earlier coverage / Tension
Both cover Models; reported by the same outlet (arxiv.org); earlier Models coverage from 2026-08-03.
Shared entity: VLMs / Same source domain / Earlier coverage / Tension
Both cover VLMs; reported by the same outlet (arxiv.org); earlier VLMs coverage from 2026-06-29.
Both cover VLMs; reported by the same outlet (arxiv.org); earlier VLMs coverage from 2026-03-20.
Shared entity: VLMs / Same source domain / Earlier coverage
Both cover VLMs; reported by the same outlet (arxiv.org); earlier VLMs coverage from 2026-08-13.