Dispatch
The M5 Ultra Mac Studio cuts time-to-first-token on a 256K prompt from 246 seconds to 104
Federico Viticci's MacStories review measures the M5 Ultra Mac Studio for local agents: 1.2 TB/s memory bandwidth (50% over the M3 Ultra's 819 GB/s), an 80-core GPU with Neural Accelerators, and 256 GB unified memory on the tested unit with a 512 GB model due in October. On Qwen3.8-Flash-Next it does roughly 2,733 tok/s prompt processing and 54-108 tok/s generation, about 2.5x the M3 Ultra on prefill, and at 256K context reaches first token in 104 seconds versus 246. An RTX 5090 still edges it on raw prompt processing at ~3,000 tok/s, but cannot hold three concurrent Flash-Next sessions the way 256 GB of unified memory can.
Source
↳ Follow the thread