Hacker News
A 17-Model, 28-Task Benchmark Run on August 23 Put GLM-5.3 First at $0.28 Per Lap Against GPT-5.5's $1.43
The Ed-o-meter, run 2026-08-23, scored 17 models across 28 tasks in coding, data, real-world, security and tool-use. GLM-5.3 cleared all five categories at 100% pass with a 9.3 rubric score, $0.28 per lap and 16.3s latency; GPT-5.5 hit 100% on security and 89% on real-world at $1.43 per lap and 13.2s; Kimi-K3 took the top rubric score at 9.5 with 96% pass but 26.4s latency. Opus 5 scored a 9.4 rubric with 100% security and real-world but had its coding runs blocked by safety filters, which is a measured instance of the refusal cost builders keep reporting anecdotally. The full suite costs about $30 to run.
↳ Follow the thread