Anthropic's Conceptual Reasoning Index Puts Opus 5 at 73.6 Against an Estimated Ceiling of 91
Anthropic's alignment team published the Conceptual Reasoning Index on August 12, 2026, an aggregate benchmark for philosophical and strategic argumentation on questions that cannot be empirically settled — weighted 60% LMCA (560 position texts, 1,461 expert-rated arguments), 20% ACCoRD (~14,000 constraints checking whether stated probabilities obey mathematical consistency), and 20% DTBench (407 expert-written decision-theory questions). As of August 10, the top model was Claude Opus 5 at 73.6 (95% CI ±2.1) against an estimated ceiling near 91. Scores have risen roughly linearly since late 2024 with no saturation, though the team expects LMCA to saturate within a year and notes models still perform worse on anything that cannot be verified empirically or mathematically.
↳ Follow the thread