Pattern: Agent Benchmark Scores Are Now Harness-Conditional — GLM-5.3's Terminal-Bench Run Used Z.ai's Own Claude Code Config and 3 Rollouts
MarkTechPost·medium signal
GLM-5.3's headline Terminal-Bench 3.0 number was produced by Z.ai running the public benchmark under its own Claude Code configuration, three rollouts per task, and generous limits — not an independent reproduction. As agent benchmarks replace single-turn evals, the harness (scaffold, tool set, retry budget, rollout count) contributes as much variance as the model, so a score without its harness spec is uninterpretable. Treat vendor agentic numbers as an upper bound and re-run on your own scaffold before switching models.