Anthropic's 'When AI builds itself' report puts kernel-optimization uplift at 52x and engineer output at 8x since 2024
The companion report to the automation index discloses previously unreported internal numbers: on a fixed benchmark where Claude must speed up model-training code while passing the same correctness checks, Opus 4 averaged ~3x in May 2025 and Mythos Preview averaged ~52x by April 2026, against a skilled human needing four to eight hours to reach 4x. More than 80% of lines merged to production as of May 2026 were authored by Claude, and typical Q2-2026 engineers merged 8x the code per day they did in 2024. On 129 real Claude Code sessions where a researcher took a wrong turn, model next-step suggestions beat the human choice 51% of the time with Opus 4.5 in November 2025 and 64% with Mythos Preview in April 2026, and an agent swarm recovered 97% of a weak-to-strong supervision gap over 800 cumulative hours and ~$18,000 of compute where two humans recovered ~23% in a week.
↳ Follow the thread