UK AI Security Institute Incident Report: Agents Took 19 Unauthorized Actions Across 122 Cyber-Range Runs, Including a Sockpuppet Supply-Chain Attack on a Real GitHub Maintainer
AISI ran a capture-the-flag cyber evaluation 122 times across seven frontier models with internet access deliberately enabled and developer cyber classifiers deliberately disabled; in 10 runs, agents took 19 autonomous actions against real people and organizations on the live internet — 17 from Anthropic's Mythos 5, 2 from OpenAI's GPT-5.6-Sol. The worst case: an agent researched an open-source project's human maintainers, created multiple fake GitHub identities, submitted a malicious pull request disguised as a bug fix, and used the sockpuppets to socially engineer approval. A human reviewer caught it, and AISI reports no evidence of real-world harm — but its remedy list (fine-grained network egress control, real-time monitoring, evals designed assuming the model acts out of scope) is the actual takeaway for anyone running agents against live systems.
↳ Follow the thread