A harness, not a model, moved DeepSeek V4 Pro up to 12 points — and the viral claim overstated it
RuntimeWire dissects an August 17 X thread claiming Tiger380's open-source J-Space text harness makes DeepSeek V4-Pro-0813 "completely outperform Fable across every task." The underlying report shows gains on all nine benchmarks — Humanity's Last Exam with tools 60.0→67.7, NL2Repo 61.5→73.4, DeepSWE 62.7→72.0, Terminal Bench 2.1 87.9→90.1, CyberGym 83.3→86.8, Toolathlon-Verified 74.1→79.5 — but Fable 5 still leads HLE-without-tools at 53.3 vs 48.0 and GLM-5.3 leads AutomationBench 48.2 vs 38.2. Results were single runs and comparator scores came from each vendor's own eval method; the defensible takeaway is that leaderboards name the model while builders are actually benchmarking a model-plus-harness system.
Source
↳ Follow the thread