Reddit
'Why Are Almost All New Benchmarks and Leaderboards Coding Focused?' Pulls 81 Comments on 51 Upvotes in r/LocalLLaMA
A r/LocalLLaMA thread questioning the coding monoculture in evaluation drew a 1.59 comment-to-score ratio (81 comments, 51 upvotes) — the highest engagement density of any technical post in today's set. The poster argues non-coding use cases go untested while coding benchmarks are also the most benchmaxxed, leaving buyers without signal on writing, extraction, or reasoning-over-documents workloads. This lands the same week APEX-Accounting showed no model clearing 2.6% pass@8 on real accounting work, which is exactly the blind spot the thread is naming.
↳ Follow the thread