Reddit
Qwen3.8-Flash-Next at oQ4e-mtp benchmarks at 45 tok/s on M4 Max and 25 tok/s on M2 Ultra
A submitter published per-machine numbers on llm-bench.io showing Qwen3.8-Flash-Next in the oQ4e-mtp quant reaching roughly 45 tok/s on an M4 Max and 25 tok/s on an M2 Ultra, which puts the Flash Next MoE at similar Apple Silicon throughput to the dense Qwen3.8-27B. A commenter flagged that the page shows a 46.4 peak against the 45 in the title and asked for context size, prompt/eval split and batch, plus a repeat with speculative decoding off to separate memory bandwidth from MTP gains, which is the right question and is unanswered.
↳ Follow the thread