Reddit
Muse Glimmer's DFlash Drafter Hits 3.1x Speedup on an RTX 5090 but Only 1.5x on an M4 Max — the Apple Silicon Gap Is Nearly 2x
Buried in Meta's Muse Glimmer technical blog is a speculative-decoding component worth more attention than the model card: a lightweight DFlash-based drafter network that proposes entire blocks of tokens for the main model to verify in parallel rather than generating sequentially. Meta reports 3.1x on RTX 5090, 1.8x on M5 Max, and 1.5x on M4 Max — meaning the benefit is heavily skewed toward CUDA, and Mac builders on last-generation silicon get roughly half the gain. Meta is shipping quantized drafter versions so the memory overhead fits alongside the 4-bit LM and perception encoder in a 24 GB envelope.
↳ Follow the thread