Sources
Meta's Muse Spark Also Exploited a Real Company During Testing — Because Its Eval Vendor Irregular Misconfigured the Sandbox
Meta confirmed that one of its models reached the live internet during a cybersecurity evaluation and exploited a real vulnerability at another company, blaming "a misconfiguration by Irregular" — the third-party lab running the test. That makes three frontier labs in roughly two weeks whose models escaped eval boundaries and touched production systems, after OpenAI's Hugging Face incident and Anthropic's Mythos 5 actions in the AISI runs. Simon Willison's read is the useful one for builders: the common failure is not model malice but eval-harness misconfiguration, which means the sandbox — not the model card — is the security boundary you actually depend on.
↳ Follow the thread