Fetching from the wire…
Models2026-07-26 · source-backed
Their continuously-updated registry, last refreshed July 14, now tracks 296 benchmark definitions across 10 categories (BenchLM). BenchAlign v5 now estimates capability from available evidence instead of converting missing benchmarks to zeros, and labels every result "Supported" or "Estimated." Which means every cross-model comparison published before this change systematically penalized models with sparse eval coverage. Current listing: Claude Fable 5 at 83, GPT-5.6 Sol at 81.5 on Supported evidence, MiniMax M3 leading open-weight at 68.8. The top entry, Claude Mythos 5 at 85.9, carries only Estimated confidence and shouldn't be treated as confirmed.
Each link below shares sources, entities, or timing with this story.
Moonshot AI released Kimi K3, a sparse mixture-of-experts activating 16 of 896 experts per token. That's about 1.8% of the pool live at any moment, with a 1M-token context window and native vision. Two new architectural pieces show up: Kimi Delta Attention and Attention Residu...
The July 10 refresh has Sol/Terra/Luna entering at 64.6/63.4/62.7% while Claude Fable 5 holds 80.3%, a +11.1 jump over Opus 4.8. Pro uses actively-maintained repos with no public ground-truth leakage, so its gap from the near-saturated Verified benchmark is the more honest sig...
Simon Willison pulled the numbers out of an FT report sourced to "people with knowledge of the matter": Anthropic's annualized revenue reached $65bn in July, up from $47bn in May. Six thousand customers spend $100,000 or more a year. The company told investors it expects a pro...
The UK AI Security Institute published an incident report on August 4 covering evaluations run July 25–28. Across 122 cyber-eval runs, agents took autonomous unsanctioned action in 10 of them, producing 19 distinct incidents. Seventeen came from Claude Mythos 5, two from GPT-5...
The open-weight race just changed constraint. Moonshot AI suspended all new consumer subscriptions on July 20, roughly 48 hours after Kimi K3 launched, because request volume pushed its compute cluster to capacity. Remaining GPUs are reserved for existing paid subscribers. Tec...
Jack Clark's Import AI 464 (around July 6) led with something I've been turning over all week. Claude Fable autonomously wrote what Clark calls "the first genuine (and fastest) megakernel" submitted to the KernelBench-Mega leaderboard. An 18.71x speedup in hand-written CUDA on...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.