Distilling Discrete Diffusion Models via Discrete MMD Unlocks Faster Inference
arXiv·medium signal
While continuous diffusion models have rich distillation methods for accelerating inference, discrete diffusion language models have lacked effective distillation techniques—until now. This paper introduces a distillation approach using Maximum Mean Discrepancy over multi-token sequences rather than single-token cross-entropy, better capturing the distributional properties of discrete generation. Opens the path to fast inference for discrete diffusion LMs, which had been stuck at slow multi-step generation.