Voices
Artificial Analysis publishes a phone-based inference benchmark: only 23 of 39 small models fit an iPhone 17 Pro at 4-bit and 8K context
The new pocket-scale leaderboard measures on real hardware, an iPhone 17 Pro and a Galaxy S26 Ultra, both 12 GB, and defines 'small' as fitting in 8 GB after 4-bit quantization including KV cache at 8K context. Only 23 of 39 candidate models cleared that bar. Results run end-to-end on a 1,024-token prompt with a 256-token response, scored across tool calling, instruction following, knowledge, scientific reasoning and quantitative reasoning, with Nanbeige, Liquid AI, Ornith AI, Alibaba and Google models on top.
↳ Follow the thread