MemToC finds instruction-tuned models keep a verified-correct answer against a wrong tool in only 6.5-17.1% of cases
MemToC, posted 26 August, is a controlled benchmark for what a tool-augmented LLM does after a tool return contradicts its parametric memory, built from 542 quality-controlled factual questions and model-specific elicited closed-book answers into 6,504 episodes with tool returns of known correctness. Across five open-weight 7-9B models, tool returns dominate: models retain a verified-correct answer against an incorrect tool in only 6.5-17.1% of eligible cases, follow a correct tool in 86.0-93.1%, and repeat the tool return in 78.4-86.0% of cases where both sources are wrong. No cross-model ordering survives three instruction-wording variants with content held fixed, which makes tool-trust behavior unstable to prompt phrasing.
Source
↳ Follow the thread