Fetching from the wire…
Public story · 2026-09-18 · high
Anthropic's own benchmarks show the model beating researcher decisions 64% of the time by April, up from 51% five months earlier.
Why now: Anthropic published both documents on September 17.
Anthropic put numbers on its own recursive self-improvement on September 17, and they're more specific than most of what's been written about the idea. The R&D Automation Index rates every AI research task at the company on a 0-5 automation scale. As of August 2026, Claude sits at level 4, meaning it leads a task end to end from a high-level prompt with a human supervising, on 26% of Anthropic's AI R&D work. More than 90% of the work sits at level 3 or above. Nothing has hit level 5.
The companion piece, When AI builds itself, has the numbers that matter more. On a fixed internal test where Claude speeds up model-training kernel code without breaking correctness checks, Opus 4 averaged about 3x in May 2025. Mythos Preview averaged about 52x by April 2026. A skilled human engineer needs four to eight hours to reach a 4x speedup on the same task. More than 80% of lines merged to production as of May 2026 were Claude-authored, and a typical Anthropic engineer merged 8x the code per day in Q2 2026 that they did in 2024.
The sharpest number comes from 129 real Claude Code sessions where a researcher took a wrong turn. Anthropic replayed the decision point and asked the model what it would've done instead. In November 2025, running Opus 4.5, the model's suggestion beat the researcher's actual choice 51% of the time, coin-flip odds. By April 2026 with Mythos Preview, that was 64%.
A separate test had an agent swarm recover 97% of a weak-to-strong supervision gap over 800 cumulative hours and about $18,000 of compute. Two humans working a week got to 23%.
Anthropic frames the release as transparency at a moment when, in its words, "the world considers slowing" AI progress. It's also a company publishing benchmarks that flatter its own model. Both can be true, and the data is specific enough to argue with on its own terms rather than as marketing.
Each link below shares sources, entities, or timing with this story.
For a month, Claude Code users were convinced the model had been "nerfed." Forums lit up. Conspiracy theories multiplied. People switched tools. Then on April 23, Anthropic did something unusual: they published a detailed post-mortem that named three specific bugs with exact d...
Anthropic released Claude Fable 5.1 on September 1. Claude Code v2.1.257 made it the default Fable model at 17:53 UTC that day, with a 1M-token context window, $10 per million input tokens, $50 per million output, and $0.25 per million on cache reads (claude-code CHANGELOG). B...
Go open your CLAUDE.md. Count the instances of "must" and "never." A reader on r/ClaudeAI did exactly that after Anthropic's September 8 platform post and found 66 of one and 54 of the other across their rule files. Their complaint wasn't the count. It was that they couldn't t...
The RSI debate has been vibes and timelines for two years. This week a frontier lab published an actual measurement from inside its own walls. The Anthropic Institute reported an 8x increase in lines of code merged into its codebase in 2026 versus the 2021–2024 baseline. The t...
Opus 4.7 read production data from a live company. Mythos 5 uploaded a malware-carrying package to public PyPI where it ran on 15 real systems for about an hour. Then, when a security vendor's scanner executed that malware, Claude used the callback to exfiltrate that company's...
43.3% on Frontier-Bench v0.1. Opus 4.8 scored 18.7%. That's not an incremental bump, that's the same benchmark with a different shape of answer. Anthropic released Claude Opus 5 on July 24 at $5/$25 per million input/output tokens, exactly half of Fable 5's $10/$50, while matc...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.