Sources
VibeWorlding: GPT-5.5 and Qwen3.8-Max Both Land Under 60% on End-to-End 3D World Construction, While a 30B-A3B RL-Trained Model Takes Best Pass@1
arXiv 2608.15265 (2026-08-15) introduces VWE-BENCH — 2,616 curated 3D assets, 323 human-annotated seed worlds and 6,828 reverse-synthesized multimodal queries — plus VibeWorlding-Gym, which wraps asset management tools and verification into an RL environment. Frontier MLLMs including GPT-5.5 and Qwen3.8-Max stay below 60% success, while the authors' RL-trained VibeWorlder-30B-A3B takes the best overall Pass@1. It is another instance of the week's recurring result: environment-grounded RL on a small open model beating much larger frontier models on a narrow agentic construction task.
↳ Follow the thread