Fetching from the wire…
Public story · 2026-07-19 · high
Each board weights cost, latency and capability differently, so a vendor's cited rank depends entirely on which one they picked.
Why now: All three trackers rolled out updates in July, making their disagreement hard to miss for anyone comparing them side by side.
Vellum, llm-stats.com and BenchLM.ai each posted refreshed July rankings for large language models. That joins HuggingFace's existing Artificial Analysis leaderboard, per Vellum's own comparison page. Four boards now, and no shared ruler between them.
That's the problem. Each board weights cost, latency and capability differently. A model could top one list and land mid-pack on another, depending on which factor gets the most weight. A vendor calling a model "ranked #1" rarely says which board, or what that board optimizes for.
For anyone picking a model, the fix is mechanical. Check two independent boards, not one. Read what each is actually measuring before trusting the number. A board built around cost weighs differently than one built around capability, even when scoring the same models.
Watch for vendors citing one favorable ranking without naming the board or its weighting. That's the tell they're shopping for a number, not reporting one.
Each link below shares sources, entities, or timing with this story.
Simon Willison released LLM / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover Artificial Analysis, July; overlapping topics (capability, july).
LLM uses OpenAI / Shared entity: July / Earlier coverage / Tension
Linked by a graph relationship (LLM uses OpenAI); both cover July; earlier July coverage from 2026-07-10.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-18.
HuggingFace released Claude Code / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (HuggingFace released Claude Code); both cover Check, LLM; overlapping topics (analysi, cost).
Simon Willison released LLM / Shared entity: July / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover July; earlier July coverage from 2026-07-14.
Linked by a graph relationship (Simon Willison released LLM); both cover July; earlier July coverage from 2026-07-10.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-19.