Research
Sycophancy Is Two Dimensions, Not One: Structured Resistance vs Compliance in LLM Moral Reasoning
Baihui Wang and Bernard Koch argue that socially calibrated models require more than minimizing sycophancy as a single scalar, and decompose model responses to social pressure into structured resistance and compliance in moral reasoning tasks. A model that never yields is as miscalibrated as one that always yields — the useful target is updating on good arguments while holding against bad ones. Practical implication: anti-sycophancy fine-tuning that optimizes a one-dimensional metric can degrade a model's ability to legitimately learn from a user.
↳ Follow the thread