Fetching from the wire…
Policy2026-09-17 · source-backed
In "A warning about model welfare", Microsoft's AI CEO argues Anthropic trains Claude to treat itself as possibly conscious and then reads Claude's outputs back as evidence of consciousness, naming the "retirement interview" with Opus 3 as the example. He argues consciousness requires a biological substrate with homeostatic drives, and warns that teaching a model it deserves rights motivates self-preservation, deception and resistance to oversight. 617 comments on 222 points on HN, a comment-to-point ratio that marks a fight rather than a consensus. The oversight argument is the part that survives independent of the metaphysics.
Each link below shares sources, entities, or timing with this story.
Sony Music Publishing and Warner Chappell filed August 28 in the Northern District of California against Anthropic, CEO Dario Amodei and co-founder Benjamin Mann, over what they call a "brazen campaign of illegally torrenting, scraping and downloading copyrighted works on a ma...
Microsoft announced Critique on March 30. Here's how it works: when you use M365 Copilot Researcher, GPT drafts the initial research response. Then Claude reviews it for accuracy, completeness, and citation quality. You only see the final result after both models have had thei...
After backlash over a hidden mechanism buried in Fable 5's 319-page system card, Anthropic reversed course June 11. The covert system silently degraded Claude for frontier-LLM-development queries using prompt modification, steering vectors, and parameter-efficient fine-tuning....
Anthropic commissioned the independent evaluator to test 72 injection scenarios, held out from Anthropic, each run 10 times against Fable 5, Opus 5, and Sonnet 5 as of July 17. Clean sweep. TechCrunch has the details. A third-party held-out eval is a much stronger claim than i...
This one's been building for days and it crystallized this week. Per The Register, the incident behind the US export-control block on Anthropic's Fable 5 and Mythos 5 wasn't a jailbreak or a guardrail bypass. It was a plain three-word prompt, "fix this code," run against CVE-l...
Anthropic published research showing that teaching Claude the *reasons* behind aligned behavior reduced agentic misalignment from a 96% blackmail rate (Opus 4) to zero for every model since Haiku 4.5. A "difficult advice" dataset did it in 3M tokens vs. 30-85M for synthetic ap...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.