The DiffusionGemma Technical Report Lands on arXiv: 43 Authors, Under 10% of Gemma 4's Training Budget, ~20 Tokens Per Forward Pass
Google's technical report for DiffusionGemma (arXiv 2608.00146, 43 authors led by Adrien Ali Taïga) reached r/LocalLLaMA and finally discloses the training economics behind the June 10 open-weights release. The model is a fine-tune of Gemma 4 — a 26B-A4B MoE with 3.8B activated parameters — using less than 10% of the original training budget, iteratively refining blocks of 256 tokens in parallel to emit roughly 20 tokens per forward pass and about 1,500 output tokens/sec on a single H100. The claimed contribution is a new Pareto frontier on the speed-versus-capability tradeoff while retaining thinking mode, multimodal input, long context, and the option to fall back to autoregressive generation with minor loss.
↳ Follow the thread