Ten Model/Harness Combinations on One Three.js Task Show a 4x Wall-Clock Spread and No Correlation With Model Tier
An independent test published 2026-09-07 ran the same prompt (a single-page Three.js sci-fi hangar with hovering drones, warning lights, emissive runway strips and volumetric fog planes) across ten model/harness pairs spanning Codex, OpenCode, OMP and DSH/PTC with GLM 5.3 Flash Max, Luna 5.6 Max, SOL 5.6 Max, Astra 6.0 Max and Qwen 3.8 27B. Qwen 3.8 27B on OpenCode finished in 8m48s while Astra 6.0 Max on Codex took 37m30s, a 4.3x spread; Luna 5.6 Max on Codex used the fewest tokens at 1,172,267 and GLM 5.3 Flash Max on OpenCode hit 96.89% cached input. The harness, not the model tier, dominated both wall-clock and cache efficiency, which is the variable most agent-tooling comparisons never isolate.
↳ Follow the thread