Fetching from the wire…
Public story · 2026-08-25 · high
The open toolkit ties two models at a 63 average score, then splits them 8 seconds and 2.3GB versus 21.4 seconds and 4GB on the same phone.
Why now: Both Pipette repos and the full result set are public as of August 25, 2026, giving on-device builders a shared benchmark to check phone-performance claims against.
Liquid AI released Pipette, an open benchmark measuring 35 model classes and 7 quantizations across quality, speed, latency and memory on real phones. The stack logged over 10,000 verified results across four devices and llama.cpp runtimes, giving app builders numbers to pick a model that fits a phone's memory and response-time budget instead of trusting vendor claims.
Under an 8GB memory and 16K context test, Nanbeige4.2-3B and LFM2.5-2.6B tied at a 63 average score. LFM2.5-2.6B answered in 8.0 seconds using 2.3GB of memory on iPhone. Nanbeige4.2-3B took 21.4 seconds and 4.0GB for the same score.
Mixture-of-experts models that activate around 1 billion parameters per token finished in under six seconds on the same test. Two repos back the project. One runs the measurement harnesses across model, quantization, runtime and device combinations. The other is a stateless, model-blind scoring service.
Model quality ties are becoming normal in these benchmarks, so speed and memory footprint, not parameter count, will decide which models ship on phones. Pipette's results and both client repos are public now, giving on-device builders a shared way to check those claims themselves.
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source domain / Shared topic / Earlier coverage
Both cover GitHub, MoE; reported by the same outlet (github.com); overlapping topics (answer, benchmark, model).
Both cover GitHub, MoE; reported by the same outlet (github.com); overlapping topics (benchmark, context, model).
Both cover GitHub, MoE; reported by the same outlet (github.com); overlapping topics (context, cover).
Both cover GitHub, MoE; reported by the same outlet (github.com); overlapping topics (against, model).
Shared entity: GitHub / Same source domain / Shared topic / Earlier coverage / Tension
Both cover GitHub; reported by the same outlet (github.com); overlapping topics (against, benchmark, model).
Both cover GitHub; reported by the same outlet (github.com); overlapping topics (against, benchmark, memory).
Both cover GitHub; reported by the same outlet (github.com); overlapping topics (against, benchmark, context).
Shared entity: GitHub / Same source domain / Shared topic / Earlier coverage
Both cover GitHub; reported by the same outlet (github.com); overlapping topics (answer, class, context, model).