Reddit
Qwen3.8-Flash-Next runs at 3.5 tok/s on a 12GB mid-range Android phone
A video post to r/LocalLLaMA shows the 80GB Qwen3.8-Flash-Next generating at 3.5 tokens per second on a 12GB mid-range Android handset, using aggressive quantization on the dense part plus unnamed optimizations, on a device the author puts at $400 to $500. The CPU hits 80C during the run, which the top commenter flags immediately, so this is a demonstration of what fits rather than something you would leave running. It is the same sparse-MoE-plus-offload trick that makes Flash-Next viable on constrained desktops, pushed to a phone.
Source
↳ Follow the thread