Dispatch
Together AI ran 900 DeepSWE rollouts and found a Flash-first cascade beats GLM-5.3 alone at under half the cost
Across 113 DeepSWE tasks with 4 trials per config (452 GLM-5.3 and 448 Flash rollouts), GLM-5.3 scored 69.0% pass@1 at $3.99 per rollout while GLM-5.3 Flash scored 63.4% at $0.24, a 17x cost gap. The gap narrows sharply with retries: 5.6 points at pass@1 collapses to 2.6 points at pass@4, which Together reads as distillation costing Flash its single-shot polish rather than its ceiling. Their recommended routing runs Flash first and escalates to full GLM-5.3 only when tests reject the answer, solving 80.9% of tasks at $1.70 each.
↳ Follow the thread