The Kuleshov Group Published a Full Build Guide for Diffusion Language Models, With Mercury 2 Claiming 1,000+ Tokens/Second
Adapted from ICLR 2026 and MLSS 2026 workshops, the guide (178 points on HN) walks masked diffusion models, block diffusion for variable-length generation, encoder-decoder architectures, remasking-based error correction, sampling distillation, guidance and RL post-training. It catalogs what is actually available to build on: LLaDA at 8B with open weights and LLaMA compatibility, Google's Gemma Diffusion, NVIDIA's Nemotron Diffusion family up to 35B, ESM3 at 100B for proteins, and Nucleotide Transformer v3 on roughly a trillion DNA tokens. The commercial datapoint is Mercury 2 at 1,000+ tokens/second on standard GPUs, which the authors put at 5-10x the throughput of comparable-quality models like Claude Haiku.
↳ Follow the thread