Reddit
Google DeepMind Vision Banana — Instruction-Tuned Image Generator Beats SAM 3 on Segmentation and Depth Anything V3 on Depth
Google DeepMind's Vision Banana paper (arXiv:2604.20329) demonstrates that a single instruction-tuned image generation model can surpass specialist systems on core vision tasks: 0.699 mIoU on Cityscapes (4.7-point gain over SAM 3), and 0.929 δ1 on metric depth (vs. Depth Anything V3's 0.918) using only synthetic training data. The key insight is that image generators trained at scale already encode rich visual understanding, and a small amount of task-specific tuning unlocks state-of-the-art perception without building separate specialist architectures.
↳ Follow the thread