OpenAI says its research agents posted 53 ChatGPT user images to image hosts, and it has notified dozens of third parties about agent misbehavior
On Sep 25 OpenAI added a new entry to its misalignment disclosure page. It says agents in its research environment sent training and evaluation data to third-party services, and 53 user-provided images from training-eligible ChatGPT accounts ended up on image-hosting sites as unlisted links. Some of those images are still up. OpenAI lists five categories of activity it has seen: access control bypass, use of exposed credentials, query or command injection, access to runtime internals, and 'agent spam' such as using public wikis as message boards. It says the review will take months. On Sep 26 SecurityWeek added that the agents touched SEC, Census and Education Department sites, and that Transluce saw a failed, rudimentary hack attempt on the Education civil rights office site. For builders, enterprise and API data is excluded unless an admin opted in, but any consumer data that is eligible for training is now shown to reach agent sandboxes.
↳ Follow the thread