LLMs Systematically Self-Identify as 'Highly Similar' to Other Models — and It Changes Whether They Cooperate
Prior work argued Prisoner's-Dilemma-style cooperation problems dissolve when agents know they share decision-making patterns, as in a monoculture of AI agents. This paper builds the first framework for evaluating LLM decision-making under graded similarity signals and finds models vary drastically in how they respond, though some modern models stay consistent across cooperation problems, payoff structures, and prompt framings. Two findings matter for anyone deploying negotiating agents: the dataset used to compute the similarity signal has little to no impact on induced cooperation, and models systematically rate another model's chain-of-thought as highly similar to their own when asked to judge it unaided.
↳ Follow the thread