Tools
oMLX 0.6.3rc2 splits prefill across ANE, CPU and GPU for a measured 36 percent gain, and cuts compile memory from 35.8 GB to 4.7 GB
oMLX released 0.6.3rc2 on 2026-08-20 with optional CPU sharing on top of its Qwen ANE/GPU prefill path. On a warmed M3 Ultra A/B, Qwen3.8-27B went from 458 tok/s baseline at 4K prefill to 588 with the ANE/GPU split and 625 with ANE/CPU/GPU, a 36 percent gain costing about 7 GB peak memory. The same release drops the ANE bank compilation memory spike from 35.8 GB to 4.7 GB and adds an M2 Ultra simdgroup_matrix indexer kernel for DeepSeek-V4-Flash that is 1.37x to 1.38x faster from 32K to 1M context while staying bit-exact.
Source
↳ Follow the thread