Tencent's Hunyuan3D-Buffalo 1.0 Unifies 3D Generation, Understanding and Editing on an 87M-Scale Multimodal Corpus, Cutting Edit Chamfer Distance 86.7%
Published August 5, 2026 (arXiv 2608.02711) and third on Hugging Face Daily Papers with 65 upvotes, Hunyuan3D-Buffalo bolts a Qwen-VL backbone onto a diffusion generator initialized from Hunyuan3D-2.1 to cover 3D understanding, text-to-3D, instruction-guided editing and text-grounded part generation in one model, trained on a purpose-built 87M-scale 3D multimodal corpus. Human evaluation gives it 56.6% overall preference against a 25% random baseline; on Edit3D-Bench it reports Chamfer Distance 0.0091 (an 86.7% reduction) and F1 0.6515 (2.39x the strongest baseline); UniPart-Bench part Q&A hits 85.47 SBERT. The paper does not state parameter counts, and neither weight availability nor a license is specified — worth checking before planning around it.
↳ Follow the thread