A Fully Fabricated Market Panel Raises LLM Commitment on Unanswerable Questions as Much as Real Data Does
arXiv 2608.27167 finds that across 12 frontier models, showing a professional-looking evidence panel drives commitment to a directional call on provably unpredictable questions from 6.5% to 54.0%, and that inventing every number on the panel still lifts commitment from 24.5% to 36.8%, statistically indistinguishable from the 37.6% produced by genuine market data. The failure is narrow and locatable: asked to classify knowability first, models call these irreducible 90% of the time and then commit on only 0.4% of those, so the act/don't-act gate is what breaks, not the belief. Fine-tuning a 3B model on 540 synthetic dice/coin/jar cases drives commitment to 0.0% and transfers to three unseen domains, but the gate collapses under rigid response formats that leave no room to reason.
↳ Follow the thread