Fetching from the wire…
Models2026-07-26 · source-backed
Their continuously-updated registry, last refreshed July 14, now tracks 296 benchmark definitions across 10 categories (BenchLM). BenchAlign v5 now estimates capability from available evidence instead of converting missing benchmarks to zeros, and labels every result "Supported" or "Estimated." Which means every cross-model comparison published before this change systematically penalized models with sparse eval coverage. Current listing: Claude Fable 5 at 83, GPT-5.6 Sol at 81.5 on Supported evidence, MiniMax M3 leading open-weight at 68.8. The top entry, Claude Mythos 5 at 85.9, carries only Estimated confidence and shouldn't be treated as confirmed.
Each link below shares sources, entities, or timing with this story.
Claude Mythos built by Anthropic / Shared entities / Earlier coverage
Linked by a graph relationship (Claude Mythos built by Anthropic); both cover Claude Fable, GPT, July, Which; earlier Claude Fable coverage from 2026-07-21.
Claude Mythos benchmarked against Claude / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Mythos benchmarked against Claude); both cover Claude Fable, GPT, July; overlapping topics (benchmark, claude).
Claude Mythos built by Anthropic / Shared entities / Earlier coverage
Linked by a graph relationship (Claude Mythos built by Anthropic); both cover GPT, July, Which; earlier GPT coverage from 2026-07-20.
Claude Fable uses CUDA / Shared entities / Earlier coverage
Linked by a graph relationship (Claude Fable uses CUDA); both cover Claude Fable, GPT, July; earlier Claude Fable coverage from 2026-07-12.
Claude Mythos built by Anthropic / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Mythos built by Anthropic); both cover Claude Fable, July; overlapping topics (claude, estimated).
Claude Mythos built by Anthropic / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Claude Mythos built by Anthropic); both cover BenchLM, GPT; reported by the same outlet (benchlm.ai).
Claude Mythos built by Anthropic / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Mythos built by Anthropic); both cover Claude Fable, July; overlapping topics (claude, confirmed).
GPT competes with Grok / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (GPT competes with Grok); both cover BenchLM, Claude Fable, GPT, July; reported by the same outlet (benchlm.ai).