Fetching from the wire…
Public story · 2026-08-07 · source-backed
The leaderboard says first place. The methodology says you should check your own bill.
Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M output. Its Intelligence Index score of 56 puts it level with Claude Opus 4.8 (max) and ahead of every model Google, Meta and xAI ship. 507 points and 320 comments on HN.
Buried in the data: it averages 64 turns on GDPval-AA. Qwen3.7 Max averaged 14. That's a ~4.5x turn-count tax, and a per-token price comparison shows you exactly none of it. If your agent loop reloads context each turn (most do), turn count is closer to a multiplier on your actual spend than a footnote.
Three other findings today say the same thing from different angles. RealReplicaBench (1,038 stars in 5 days) ran 12 models across 107 long-horizon tasks in high-fidelity stateful clones of eight commerce and logistics platforms. Claude Opus 5 led at 66/107 (61.7%) on the Accio harness and 60/107 (56.1%) on OpenClaw. Same model, same tasks, 5+ points of swing from swapping the scaffold. DCAS found that open coding models fine-tuned on OpenHands trajectories degrade substantially under any other CLI scaffold, while untrained base models show no such divergence, pinning the load-bearing variable on planning structure. And the single largest sentiment signal across tracked subreddits today was a meme mocking benchmark charts at 5,938 upvotes with only 56 comments. Low comment ratio means consensus, not argument. Nobody's defending vendor evals.
The arithmetic to run before you swap models: measure average turns per completed task on your workload, multiply by your average context size, multiply by input price. Then compare. I'd bet on the cheaper-per-token model losing that comparison more often than the leaderboards imply, and I'd also bet most teams have never run it.
Related and worth pricing in: DeepSeek emailed API users on August 6 warning of a "significant" price increase, citing demand beyond platform capacity and unsustainable server costs. V4-Flash currently sits at $0.14/M input and $0.28/M output, the floor anchoring the whole cheap-inference market. Second pricing change in under a month after the mid-July peak/off-peak split. Continued use after adjustment constitutes acceptance. If your cost model has a DeepSeek number in it, that number is stale.
Each link below shares sources, entities, or timing with this story.
OpenHands uses GPT / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (OpenHands uses GPT); both cover August, Claude Opus, DeepSeek, Flash; reported by the same outlet (github.com, reddit.com).
Gemini competes with ChatGPT / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Gemini competes with ChatGPT); both cover Artificial Analysis, DeepSeek, Flash, GPT; reported by the same outlet (artificialanalysis.ai).
Claude competes with ChatGPT / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Claude competes with ChatGPT); both cover Claude Opus, DeepSeek, Flash, Then; reported by the same outlet (reddit.com).
Gemini competes with ChatGPT / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Gemini competes with ChatGPT); both cover Buried, Claude Opus, Flash, GPT; reported by the same outlet (arxiv.org).
OpenHands uses GPT / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (OpenHands uses GPT); both cover DeepSeek, Fable, Flash, GPT; overlapping topics (agentic, context, cost, model).
OpenHands uses Python / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (OpenHands uses Python); both cover Claude Opus, Fable, GPT, Same; reported by the same outlet (arxiv.org).
Claude competes with ChatGPT / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Claude competes with ChatGPT); both cover GPT, July, Nobody, Same; reported by the same outlet (arxiv.org).
ChatGPT built by OpenAI / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (ChatGPT built by OpenAI); both cover Claude Opus, DeepSeek, Google, GPT; overlapping topics (agentic, context, cost, model).