Reddit
First M5 Ultra inference numbers surface on omlx.ai: 50 tok/s generation and 1,800 tok/s prefill on Qwen3.8 27B q4
An r/LocalLLaMA post surfaced what appear to be the first M5 Ultra entries on the omlx.ai benchmark database: for Qwen3.8 27B at q4 with an 8K context and no MTP, 50 tok/s token generation and 1,800 tok/s prompt processing. The thread's value is the pushback rather than the headline: an M3 Ultra 512GB owner reports 20-25 tok/s on q8 of the same model, meaning the generation gain is modest and the real jump is prefill, and the same commenter notes multiple entries for identical hardware and model configs varying by more than 10 tok/s. Treat the table as crowd-submitted and unofficial, not an Apple datasheet.
↳ Follow the thread