Reddit
A 3.5-hour, 100-turn run of Qwen3.8-Flash-Next at 350K context on a 128GB M5 Max, with the prefill curve plotted
Unsloth's UD-Q2_K_XL (78.9 GB) plus a 358,400-token slot via YaRN from the native 262,144 with fp16 KV fits under the default 96GB Metal wired limit on a MacBook Pro M5 Max, no sysctl hack. Cold prefill runs 1,561 t/s at 5.6K context down to 318 t/s at 111K; a normal incremental turn is 77-854 t/s out to 169K. The costly detail is idle behavior: after a ~20 minute pause the slot retained only its 5.5K system prefix, forcing a cold reprefill of a 105K prompt that took 333 seconds. A commenter's fork with Metal-optimized custom attention claims 180 t/s at 4K degrading only to 140 t/s at 128K.
Source
↳ Follow the thread