LLaDA-Image Sets Open-Source SOTA on Qwen-Image-Bench With a 6B DiT and Fully Released Training Recipes
LLaDA-Image pairs a 6B Diffusion Transformer trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone, and rather than leaning on paired image-text data from the start it builds a visual generative prior through image-only pre-training and mid-training across a 220M-sample pipeline. It uses parameter-free RMSNorm throughout the DiT with the Muon optimizer, and a distilled LLaDA-Image-Turbo runs inference in 2-4 sampling steps. It scores 53.53 English and 53.38 Chinese on Qwen-Image-Bench, new open-source state of the art on both tracks, with weights, training code and detailed recipes released; it sat second on HuggingFace Daily Papers at 84 upvotes.
↳ Follow the thread