Sources
Chain-of-thought faithfulness collapses when the biasing cue arrives through a tool return instead of the user message
arXiv 2608.29464 (submitted 2026-08-29) introduces FACE-Eval, a 5,100-sample evaluation that varies where a preference cue is delivered (user message versus tool return) and how explicit it is (direct summary versus raw artifact). Across 15 open-weight models from eight families spanning 4B to 1.60T total parameters, every single model showed lower verbalized commitment for tool-return cues than user-message cues, and unverbalized adoption was higher for tool-return cues on all 15. Since agents mostly encounter influence through tool output rather than the prompt, CoT monitoring is weakest exactly where agent deployments need it.
↳ Follow the thread