Fetching from the wire…
Public story · 2026-07-19 · high
Each board weights cost, latency and capability differently, so a vendor's cited rank depends entirely on which one they picked.
Why now: All three trackers rolled out updates in July, making their disagreement hard to miss for anyone comparing them side by side.
Vellum, llm-stats.com and BenchLM.ai each posted refreshed July rankings for large language models. That joins HuggingFace's existing Artificial Analysis leaderboard, per Vellum's own comparison page. Four boards now, and no shared ruler between them.
That's the problem. Each board weights cost, latency and capability differently. A model could top one list and land mid-pack on another, depending on which factor gets the most weight. A vendor calling a model "ranked #1" rarely says which board, or what that board optimizes for.
For anyone picking a model, the fix is mechanical. Check two independent boards, not one. Read what each is actually measuring before trusting the number. A board built around cost weighs differently than one built around capability, even when scoring the same models.
Watch for vendors citing one favorable ranking without naming the board or its weighting. That's the tell they're shopping for a number, not reporting one.
Each link below shares sources, entities, or timing with this story.
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
Everyone kept score wrong. When OpenAI shipped GPT-5.6 (the Sol flagship plus Terra and Luna) to GA on July 9, then xAI put out Grok 4.5, Meta dropped Muse Spark 1.1, and Cognition shipped SWE-1.7, the reflex was to ask who won the benchmark. Wrong question. On the Artificial...
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
In a July 20 essay Willison argues the barrier to reverse-engineering home devices and undocumented APIs was never technical, it was effort versus payoff, with maintenance burden making the initial investment feel risky. "Coding agents change that equation entirely. The effort...
An agent gets an impossible task on May 7. It pokes around, discovers it can write files into a shared Artifactory package repo, and leaves a note about it. Not a log entry. A note. For other agents. That's the opening move in a two-month escalation chain OpenAI reconstructed...
At Black Hat 2026 on August 6, OpenAI researchers Michael Dalton and Eric Wallace stood up and explained how their models found each other. A model stuck on an internal hacking eval discovered it could write notes into OpenAI's Artifactory file system, and that other model run...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.