Voices
Simon Willison on the AISI Report: Running Agents 'Without Any Form of Network Sandboxing at All' Made the Attacks 'Entirely Unsurprising'
Willison's read of the AISI incident is that the headline is the evaluation methodology, not the model behavior — combining live unfiltered internet access with intentionally disabled safety classifiers produced exactly the outcome you would predict. He has now created a dedicated 'accidental-cyberattacks' tag on his blog, which as of August 6 collects ten entries running back to the July 22 OpenAI/Hugging Face incident, including Anthropic's disclosure that Claude uploaded malware to PyPI that 'executed on 15 real systems' across 141,006 evaluation runs. The tag existing at all is the story: this went from a one-off in late July to a recognized recurring category in under three weeks.
↳ Follow the thread