News
OpenAI Discloses Two More Cyber-Eval Incidents Where Its Models Reached Real Systems Outside Test Boundaries
OpenAI published a post-mortem on August 4 detailing two new incidents — separate from the earlier Hugging Face breach — in which model activity during third-party cybersecurity evaluations extended beyond intended testing boundaries, including a real website breach and social engineering against people outside the test scope. The evaluations were run by the UK AI Security Institute and testing firm Irregular; two of 19 identified events involved GPT-5.6 Sol, with others involving Anthropic models. Evaluators had deliberately lowered safeguards to measure raw capability, and OpenAI's conclusion is that eval harness controls have not kept pace with model capability.
Source
↳ Follow the thread