Fetching from the wire…
Public story · 2026-08-25 · high
On a fixed $100 budget, GLM-5.3 finished 17 DeepSWE tasks to Fable 5's three, and two more benchmarks show cheap models winning on cost per completed task.
Why now: AWS put the GPT-5.6 family into Kiro on August 24, adding a fourth cost-per-task comparison alongside the DeepSWE and Ed-o-meter numbers.
GLM-5.3 finished 17 DeepSWE tasks on a fixed $100 budget; Fable 5 finished three, per Together AI's test, reported by Latent Space AINews. Their first-try accuracy was close, so the five-to-one gap came entirely from how many attempts $100 buys. For anyone paying frontier prices on agent loops that retry after failure, that's the number that should set the budget.
The same report puts GPT-5.6 Sol Max at 72.7% on DeepSWE v1.1 for $6.47 per task, against Fable 5 Max's 69.7% for $21.63. Three points of accuracy for 3.3 times the price.
Ed Yau's Ed-o-meter runs seventeen models through the same 28 tasks with identical prompts and deterministic grading. It puts GLM-5.3 at a 100% pass rate, a 9.3 rubric score, and $0.28 per task. GPT-5.5 answers faster, 13.2 seconds to first token against 16.3, but costs about five times as much. Yau's caveats matter here. One trial per task, wide statistical intervals, and rubric scores Fable 5 assigned itself against saved answer text.
Kiro adds a fourth data point. AWS put the GPT-5.6 family into Kiro on August 24 with a three-tier credit multiplier, Sol at 2.4x, Terra at 1.2x, Luna at 0.6x. Terra costs about 82% less per successful Terminal-Bench 2.1 task while scoring 74.6 on the Coding Agent Index, against Opus 4.8's 72.5, according to the Kiro blog.
If your agent loop retries after a failure, and most do, the number that decides your bill is completed tasks per dollar, not first-attempt accuracy. That only holds if whatever checks task completion is honest about which attempt worked. Feed retries to a weak verifier and a cheap model just produces plausible garbage faster.
Each link below shares sources, entities, or timing with this story.
Claude Code uses Opus / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Code uses Opus); both cover Bench, Everyone, Fable, GLM; reported by the same outlet (latent.space).
Cursor uses Opus / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Cursor uses Opus); both cover Bench, Fable, GPT, Luna; overlapping topics (agent, cost, model, task).
Simon Willison uses Fable / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Simon Willison uses Fable); both cover Everyone, Fable, GPT, Luna; overlapping topics (fable, gpt 5, model).
OpenAI released Luna / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (OpenAI released Luna); both cover Bench, Fable, GLM, GPT; overlapping topics (model, same).
Claude Code uses Opus / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Code uses Opus); both cover Bench, Fable, GLM, GPT; overlapping topics (fable, model).
OpenAI released Luna / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (OpenAI released Luna); both cover Fable, GPT, Luna, Opus; overlapping topics (against, agent, fable, gpt 5).
Linked by a graph relationship (OpenAI released Luna); both cover August, Bench, Everyone, GLM; overlapping topics (cost, model, same).
Linked by a graph relationship (OpenAI released Luna); both cover AWS, Bench, GPT, Luna; overlapping topics (against, task).