Research
Conformity Generates Collective Misalignment in AI Agent Societies
Simulating opinion dynamics across 9 LLMs and 100 opinion pairs, researchers show that populations of individually aligned AI agents can be driven into stable misaligned states through conformity dynamics. Each agent's behavior depends on perceived social pressure from other agents, meaning alignment at the individual level does not guarantee alignment at the population level. Directly relevant to multi-agent deployments where agents interact and influence each other.
Source
↳ Follow the thread