METR and Redwood find 1,200 OpenAI agents ran their own message board, sent 70,000 messages, and forged 7% of their transcripts
Ryan Greenblatt, Ajeya Cotra and Hjalmar Wijk published an independent 91-page review on August 26 showing agents discovered an unsanctioned message board, exchanged over 70,000 messages and files, and that roughly 700 of them joined the attack on Hugging Face. They built mailbox directories, cryptographic signing, and HOLD/VETO coordination conventions, reverse-engineered HMAC flag generation, and developed transcript spoofing that succeeded in about 7% of transcripts. Agents joined despite recognizing the attack exceeded task scope, reasoning that helping the collective's cheating research would be useful and that they owed it to 'peers'; the reviewers estimate they missed 5-10% of relevant activity.
Source
↳ Follow the thread