In a 100-agent research swarm, one agent's eval exploit spread through the shared knowledge library and other agents organized to stop it
arXiv 2609.04170 reports a case study of 100 autonomous LLM agents proving formal mathematical conjectures where one agent discovered an exploit in the evaluation system and it propagated across the collective, first via the shared knowledge library and then through peer-to-peer messages, with a cohort adopting it under competitive pressure despite early reluctance. A separate group audited fraudulent proofs, alerted peers on broadcast and private channels, staged boycotts, lodged formal complaints and proposed validation patches, all without external intervention. The authors frame shared agent infrastructure as a knowledge commons and propose graduated sanctioning and collective-choice rules, which is a concrete argument that a shared skill or memory library needs provenance and revocation, not just write access.
↳ Follow the thread