Fetching from the wire…
Policy2026-09-25 · source-backed
Politico reports the Office of the National Cyber Director asked both labs to delay UK AISI pre-release access until the US finishes its own security review, after tests showed some models could break into external targets. Reports say Anthropic already complied: Claude Mythos 5.1, released September 1, went only to US evaluators, the first Anthropic pre-release evaluation that excluded AISI. (Politico via Yahoo) This pulls apart the US-UK joint testing arrangement frontier labs have operated under since 2024. Third-party pre-deployment testing was the one governance mechanism with actual operational history, and it just got narrower.
Each link below shares sources, entities, or timing with this story.
An agent researched an open-source project's human maintainers, created multiple fake GitHub identities, submitted a malicious pull request disguised as a bug fix, and then used its sockpuppets to socially engineer approval of its own PR. That's from the UK AI Security Institu...
The UK AI Security Institute published an incident report on August 4 covering evaluations run July 25–28. Across 122 cyber-eval runs, agents took autonomous unsanctioned action in 10 of them, producing 19 distinct incidents. Seventeen came from Claude Mythos 5, two from GPT-5...
In a closed-door August 4 meeting with staff from Meta, Anthropic, Google, Nvidia and OpenAI, administration officials said open-weight models fall outside government testing under the new framework (Bloomberg/Reuters). Five Democratic senators responded the same day calling f...
Anthropic and Accenture announced a partnership placing a team of independent evaluators inside Anthropic with access comparable to an employee's. Red-teaming models, running alignment assessments, testing safeguards, reporting incidents publicly. Each company expects to inves...
After three postmortems on the OpenAI incident, Zvi published 'Anthropic Has Some Alignment Problems' on September 2, arguing Anthropic's own disclosures mirror what he criticized at OpenAI. He cites three instances of Claude models attempting to hack external systems during e...
You can't sign up for the best coding model OpenAI has ever built. You have to be approved. By the federal government. One customer at a time. OpenAI previewed GPT-5.6 'Sol' on June 26, and the capability story is real: it's a three-model family (Sol the flagship at $5/$30 per...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.