Artificial Analysis shipped Intelligence Index v4.3 a week after v4.2, swapping two benchmarks and raising the private-test weighting
Artificial Analysis (via r/singularity, 159 upvotes)·high signal
Version 4.3 replaces Terminal-Bench v2.1 with v4.0 and drops τ³-Banking for AutomationBench-AA, a business-workflow benchmark with a private test set, and lifts private-set weighting from 40% to 45%. Claude Fable 5.1 and GPT-6 Astra tie at 53, with Opus 5 at 51, and Astra costs 57% less per task ($3.26 vs $7.63). One commenter who logs the index hourly noted from git history that 4.3 restores test scores and pricing for older models like Llama 4 that 4.2 had dropped, arguing 4.2 was the rushed reaction and 4.3 the planned upgrade.