Bigger models trust stale memory harder, and hiding the label amplifies it at every size
The Memory Trust Gap benchmark uses two suites on a same-family size series (Qwen3 0.6/1.7/4/8B): a Benefit suite unsolvable without the stored fact, and a Safety suite where an authoritative tool always holds the correct value. Models answer with the stale stored value 0.92 to 1.00 of the time at every scale in the Benefit suite, and in the Safety suite the harm below the no-memory baseline is capability-gated, with larger models collapsing hardest once a stale note is made to look current. A 2x2x2x2 factorial shows removing a label amplifies over-trust at every size while a recency feature fools larger models more, and source authority is weak and scale-flat. This is over-trust, not confusion, and mitigation is itself capability-dependent.
Source
↳ Follow the thread