Fetching from the wire…
Models2026-08-27 · source-backed
Three points behind GLM-5.3 at 60, tying GPT-5.6 Terra and Muse Spark 1.2, at $0.09 per task against $0.68 for GLM-5.3 max (Latent Space). It burned 149M output tokens to run the index, of which 134M were reasoning tokens, more than Kimi K3 at 133M or Qwen3.8 2.4T A95B at 136M at comparable scores. The economics come from $0.15/$0.50 per million in and out, not token frugality. Budget in dollars, not tokens. Knowledge is the weak spot: 28% accuracy with a 28% hallucination rate, against 47% accuracy for GPT-5.6 Terra, while Terminal-Bench v2.1 reaches 84.3%.
Each link below shares sources, entities, or timing with this story.
OpenAI released Terra / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (OpenAI released Terra); both cover Bench, Flash, Knowledge, Muse Spark; overlapping topics (against, token).
OpenAI released Terra / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (OpenAI released Terra); both cover Bench, Flash, GPT, Terminal; overlapping topics (gpt-5, reasoning, task, token).
Kimi K3 benchmarked against Fable / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Kimi K3 benchmarked against Fable); both cover Artificial Analysis, Bench, GPT, Terminal; overlapping topics (against, analysi, artificial, behind, gpt-5).
Linked by a graph relationship (Kimi K3 benchmarked against Fable); both cover Bench, GLM, GPT, Terminal; overlapping topics (behind, glm-5, gpt-5).
OpenAI released Terra / Shared entities / Earlier coverage
Linked by a graph relationship (OpenAI released Terra); both cover Bench, GLM, GPT, Terminal; earlier Bench coverage from 2026-06-26.
Kimi K3 benchmarked against Fable / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Kimi K3 benchmarked against Fable); both cover Artificial Analysis, Bench, GPT, Terminal; overlapping topics (gpt-5, output, task, token).
Kimi K3 benchmarked against Fable / Shared entities / Earlier coverage
Linked by a graph relationship (Kimi K3 benchmarked against Fable); both cover Artificial Analysis, GPT, Kimi K3, Muse Spark; earlier Artificial Analysis coverage from 2026-07-17.
Kimi K3 benchmarked against Fable / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Kimi K3 benchmarked against Fable); both cover Artificial Analysis, Flash, GPT, Qwen3; overlapping topics (index, task).