VLMs Know When They Should Abstain — Probes Hit 0.91 AUROC While Best Spontaneous Restraint Scores 0.292
TRAPSBench (arXiv 2608.13167, 2026-08-13) is a procedurally generated video benchmark of 1,404 matched physics pairs where a single targeted change makes the outcome undeterminable from the visuals, scored by a new Penalized Epistemic Calibration Score that demands both correct answers when knowable and abstention when not. Across 16 VLMs in five families the best PECS is 0.292, but linear probes decode answerability from hidden states at up to 0.91 AUROC and steering a single-layer 'void' direction causally induces or suppresses abstention. Models detect textual impossibility about 4x more readily than missing visual evidence; the authors conclude the fix must be an output-stage intervention, not better perception.
↳ Follow the thread