Sources
Under a Fixed $100 Budget on DeepSWE, GLM-5.3 Solved 17 Tasks to Fable 5's 3 Despite Similar First-Try Accuracy
Together AI's cost-normalized run, summarized in AINews on 2026-08-25, gave both models the same $100 and measured completed work rather than pass rate, finding GLM-5.3 finished roughly five times as many DeepSWE tasks. A second data point from @reach_vb puts GPT-5.6 Sol Max at 72.7% on DeepSWE v1.1 for $6.47 per task against Fable 5 Max at 69.7% for $21.63. The pattern to take from this is that per-task accuracy and per-dollar throughput now rank models differently, so an agent loop that retries cheaply beats a stronger model you can only afford to run a few times.
↳ Follow the thread