Voices
Artificial Analysis moves 40% of its Intelligence Index to private held-out sets, double v4.1
The v4.2 update published September 4 adds AA-Briefcase, a private eval of agentic knowledge work on expert-built projects, and GDP.pdf, a document-reasoning test over 4,592 pages of tables, charts and footnotes. GPQA Diamond was dropped as saturated. Claude Fable 5.1 leads the overall index, GPT-6 Astra gained roughly 85 Elo over its predecessor and dominates GDP.pdf at 33.2%, and Meta ranks third among labs. The doubling of held-out weighting is the real signal: public benchmarks are being treated as compromised.
↳ Follow the thread