Qwen ships Qwen-Image-2.1: a 7B visual generator that tops the open leaderboard and generates real RGBA transparency
Alibaba's Qwen team open-sourced Qwen-Image-2.1 on 2026-09-20, unifying text-to-image, image editing and native transparent-image generation in one model whose visual generator is 7B parameters across 32 single-stream DiT layers, paired with a Qwen3-VL 8B text encoder and a 64-channel RGBA VAE at 16x spatial compression. It scores 60.28 on the public open-source leaderboard, ahead of Nano Banana 2.0 (59.82) and GPT Image 1.5 (59.65), accepts up to 10 reference images, supports bounding-box/brush/mask local edits and native 2K output. Weights landed simultaneously on Hugging Face, ModelScope and GitHub, and the visual generator is down from roughly 20B in the original Qwen-Image, so a builder can run frontier-class editing on far less VRAM than the closed competitors it beats.
↳ Follow the thread