Artificial Analysis pushed Intelligence Index v4.2 with private held-out test sets and dropped IFBench as saturated
Artificial Analysis·medium signal
Announced 2026-09-04 and at 156 points on HN, v4.2 is an interim release pulling forward parts of v5. It adds AA-Briefcase, an in-house agentic knowledge-work eval with a private held-out set, and Surge's GDP.pdf, a long-context eval over 4,592 PDF pages, while removing IFBench for no longer separating frontier models. The ten-eval composite now puts Claude Fable 5.1 at 57, GPT-6 Astra (max) at 55 and GPT-6 Astra (xhigh) at 54.