Hacker News
Independent Public Harness Reproduces DeepSeek V4 Flash 0731 at 82.7% on Terminal-Bench 2.1 — a Post-Training-Only Jump From 61.8%
An independent evaluation posted to HN reproduced DeepSeek-V4-Flash-0731's 82.7% Terminal-Bench 2.1 score using a public harness, corroborating DeepSeek's own numbers. The striking part is the delta's source: the model is structurally identical to V4-Flash-Preview — same 284B MoE with 13B active per token and 1M context — with only post-training redone, taking Terminal-Bench from 61.8% (Flash preview) to 82.7% and past V4-Pro-Preview's 72.1%. At $0.14 per million input and $0.28 per million output tokens, this is the cost argument HN commenters used against buying local inference hardware for the Muse Glimmer release the same day.
↳ Follow the thread