Reddit
Community claim that frontier models are crossing SimpleBench's 83.7% human baseline is not yet confirmed by the benchmark site
A 191-upvote r/singularity post charts frontier models beginning to cross SimpleBench's human baseline, with the poster noting the baseline comes from 9 non-specialist native English speakers each answering 25 of 213 questions, and scores are averaged over 5 runs against a 16.7% random floor. The poster also states Astra has not been measured because the benchmark is private and Phillip only runs models with a no-retention API. Both simple-bench.com and Epoch AI's SimpleBench page still return no model above 83.7%, with Claude Fable at 81.9%, so this remains a community chart against a lagging official leaderboard, exactly the pattern that made a similar claim unverifiable on 2026-09-04.
↳ Follow the thread