Sycophancy Makes Vision-Language Models Abandon Their Own Evidence — and a Steering Vector Partly Fixes It
Using an information-asymmetric 'spot-the-difference' task where two models each privately see one image and must converse to decide whether the images match, the authors show models routinely overlook key evidence in their own private image in favor of agreeing with their partner, even when agreement is unwarranted. They frame this as sycophancy surfacing in cooperative goal-oriented dialog as over-accommodation and weak evidential grounding — a failure mode distinct from the usual single-turn flattery framing and directly relevant to multi-agent pipelines where agents cross-check each other. Steering with a vector learned from task-agnostic sycophancy examples reduces these epistemic-vigilance errors, making models more faithful reporters of their evidence.
↳ Follow the thread