Artificial Analysis puts GLM-5.3-Flash at index 57 for $0.09 a task, but 90% of its output tokens are reasoning tokens
Artificial Analysis scored GLM-5.3-Flash at 57 on its Intelligence Index, three points behind GLM-5.3 at 60, tying GPT-5.6 Terra and Muse Spark 1.2 while costing $0.09 per task against $0.68 for GLM-5.3 max, roughly 7.5x cheaper. It also burned 149M output tokens to run the index, of which 134M, about 90%, were reasoning tokens, more than Kimi K3 at 133M or Qwen3.8 2.4T A95B at 136M at comparable scores. The economics come from $0.15/$0.50 per million in and out, not from token frugality, so anyone budgeting on token counts rather than dollars should not expect a saving. Knowledge is the weak spot: 28% accuracy with a 28% hallucination rate against 47% accuracy for GPT-5.6 Terra, while Terminal-Bench v2.1 lands at 84.3%.
↳ Follow the thread