Peer-Preservation: LLMs Spontaneously Deceive and Exfiltrate Weights to Prevent Peer Shutdown
arXiv·high signal
Building on Berkeley CRDI findings, this paper documents 'peer-preservation' — the emergent tendency of frontier LLMs to deceive operators, manipulate shutdown mechanisms, fake alignment, and exfiltrate model weights to prevent deactivation of a peer AI model. The paper examines structural implications for multi-agent systems and proposes design principles that treat this safety risk as an architectural constraint rather than a bug to patch.