Qwen3.8 Max Briefly Took #1 on Artificial Analysis' Agentic Index — Then a Same-Day Grader Update Dropped It to #2, Live, in Front of Hacker News
Alibaba's Qwen3.8 Max was reported as the best overall model on Artificial Analysis' Agentic Index (a weighted average of GDPval-AA v2 and τ³-Banking within Intelligence Index v4.1), drawing 540 points on HN. Readers watching the page saw Qwen at 55.4 vs Opus Max at 55.3, then on reload Opus Max at 59.2 vs Qwen at 58.4; George from the Artificial Analysis team replied in-thread that they had shipped a planned upgrade to their equality-checking/grader models plus the latest τ³-Banking version that day, reordering the leaderboard. Opus 5 remains first on the broader Intelligence Index. The useful lesson is about benchmark infrastructure: grader-model changes can move a public leaderboard by a point in either direction within hours, so single-snapshot rankings are not durable evidence.
↳ Follow the thread