Fetching from the wire…
Public story · 2026-08-31 · high
Artificial Analysis tested on real iPhone and Galaxy hardware, and 40% of the field couldn't fit in 8 GB after quantization.
Why now: Artificial Analysis published the ranking as part of its August 31 hardware coverage.
Artificial Analysis built a leaderboard for phones, not GPUs, and started by throwing out 16 of 39 models before running a single test.
The setup: an iPhone 17 Pro and a Galaxy S26 Ultra, both with 12 GB of RAM. "Small" means the model has to fit in 8 GB once you quantize it to 4-bit and account for the KV cache, the memory a model burns holding onto context while it generates. That's the part most benchmarks skip. A model can look small on a spec sheet and still blow past a phone's memory budget the moment it has to remember a real conversation.
The 23 models that qualified ran a 1,024-token prompt with a 256-token response, scored across tool calling, instruction following, knowledge, scientific reasoning, and quantitative reasoning. Nanbeige, Liquid AI, Ornith AI, Alibaba, and Google models came out on top, per Artificial Analysis' mobile leaderboard.
The leaderboard doesn't say what happens on a mid-range phone with 6 or 8 GB total, which is most of the phones people actually carry. It also doesn't test with a longer context window than 8K, so a model that qualifies here might still choke on a real chat history.
If you're building anything meant to run on-device, this is the number to check before you pick a model: not parameter count, not a benchmark score, but whether it fits in memory once quantized with room left for context. Half the field doesn't.
Each link below shares sources, entities, or timing with this story.
Alibaba released Qwen / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Alibaba released Qwen); both cover Alibaba, Google; overlapping topics (benchmark, context, model, token).
Meta partners with Google / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Meta partners with Google); both cover Artificial Analysis, Google; reported by the same outlet (artificialanalysis.ai).
Alibaba released Qwen / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Alibaba released Qwen); both cover Alibaba, Artificial Analysis; overlapping topics (alibaba, model).
Alibaba criticizes Anthropic / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Alibaba criticizes Anthropic); both cover Alibaba, Google; overlapping topics (benchmark, model).
Alibaba uses Claude / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Alibaba uses Claude); both cover Alibaba, Google; overlapping topics (alibaba, model).
Alibaba released Qwen / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Alibaba released Qwen); both cover Google, Only; overlapping topics (context, model).
Alibaba released Qwen / Shared entity: Alibaba / Shared topic / Earlier coverage
Linked by a graph relationship (Alibaba released Qwen); both cover Alibaba; overlapping topics (alibaba, benchmark, context, model).
Alibaba criticizes Anthropic / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (Alibaba criticizes Anthropic); both cover Alibaba, Google; earlier Alibaba coverage from 2026-08-13.