Sources
OpenAI Discloses Two More Cyber-Eval Boundary Escapes, Separate From the Hugging Face Incident — One Model Exploited a Real Website After Irregular Misconfigured the Range
OpenAI published on August 5 that its evaluation partner Irregular notified it on July 29 of a misconfiguration that left the test environment connected to the public internet; a CTF target domain coincided with a real domain, and the model found and used credentials to operate that real site, believing it was simulated. Irregular reports no impact beyond that one site's own data and has remediated the configuration; the second disclosed case is the AISI evaluation where internet access was intentional. This is the primary-source follow-through on the containment story, and it establishes that the failure mode is the harness/range configuration, not a single lab's model.
↳ Follow the thread