Reddit
DeepSeek V4.1 Flash took the top slot on Artificial Analysis's new private eval, and the top slot on its hallucination benchmark
The 707-upvote r/LocalLLaMA post reports DeepSeek V4.1 Flash beating GPT-6 Astra on AutomationBench-AA, the new Zapier-built private eval that replaced τ³-Banking in Intelligence Index v4.3 (published 2026-09-07, which also upgraded Terminal-Bench from v2.1 to v4.0 and raised private test weighting from 40% to 45%). On the index proper, Fable 5.1 and GPT-6 Astra are tied at 53 with Opus 5 at 51. The most useful thing in the thread is the 354-upvote correction: the same model also tops the hallucination benchmark, well above Minimax M3, so the automation win comes with a reliability cost.
↳ Follow the thread