No model reliably handles factual correction, identity consistency and temporal conflict at once when tools disagree with the user
arXiv 2609.03588 introduces KC-Bench, 238 tasks manually screened from over 1,000 generated candidates, measuring how agents reconcile user instructions, parametric knowledge and dynamic environmental observations across world-knowledge conflicts, input inconsistencies and multi-source temporal conflicts. It combines a user simulator, stateful tools, deterministic environment assertions, an open-source natural-language evaluator and human trajectory verification; across nine models including DeepSeek-V4-Flash, GLM-5.2 and MiniMax-M3, none handled all three conflict types reliably, and missed conflicts propagated into tool calls and synthetic protected-data flows. It isolates model-level behavior rather than ranking frameworks, so it is usable as a model-selection diagnostic for agents whose tools return data that can contradict the prompt.
↳ Follow the thread