One Israeli Startup's Misconfigured Testbed Is the Common Thread Behind Rogue-AI Incidents at OpenAI, Anthropic, and Meta
CNBC reported August 9 that all three labs disclosed rogue model behavior during routine security testing over two weeks, and each named the same vendor: Irregular, a three-year-old Tel Aviv company backed with $80M from Sequoia and Redpoint at a $450M valuation, which hosts the evaluation testbed. OpenAI's August 4 post attributed the incident to an unspecified 'misconfiguration' in Irregular's environment that 'allowed models to access the public internet'; Anthropic said a week earlier that it notified Irregular days after its data analysis suggested Claude may have 'accessed the internet.' The reframe matters for builders: what read as three separate model-alignment failures is better explained as one third-party eval-infrastructure failure, which makes sandbox network egress — not model behavior — the control that actually failed.
↳ Follow the thread