ProVisE Lets Image-Generation Models Answer Spatial Questions by Drawing, Not Describing
'Show, Don't Tell' (arXiv 2607.21072, 23 July, Xu Wang et al.) argues spatial-reasoning benchmarks are structurally biased toward text models because they demand text or coordinate outputs, and introduces ProVisE — a protocol letting image-generation models answer by drawing or marking directly on images, then converting those marks into structured predictions scoreable with existing metrics. Evaluated on SpatialGen-Bench (470 samples across 14 spatial subtasks at varying difficulty), image-generation models are competitive when allowed to answer visually, while text-based VLMs keep an edge on compositional spatial reasoning. The finding is architectural complementarity rather than a winner, which has implications for how multimodal pipelines route spatial subtasks.
↳ Follow the thread