A Unified Multimodal Model for Scalable 3D Generation,
Understanding, and Editing
Junliang Ye*, Kenkun Liu*, Guocun Wang*,
Yang Li*†, Yansong Qu*, Chunshi Wang*,
Jingwei Xu, Yunhan Yang, Zibo Zhao, Jiachen Xu,
Jiaao Yu, Lifu Wang,
Zhihao Liang, Zhuo Chen†, Chunchao Guo†
Tencent Hunyuan3D
* Core Contributors
3D Editing
Junliang Ye, Guocun Wang, Yansong Qu, Yang Li, Chunshi Wang, Kenkun Liu
Text-to-3D
3D Understanding
Guocun Wang, Junliang Ye, Kenkun Liu, Yang Li
A unified multimodal model for 3D generation, understanding, and editing.
Hunyuan3D-Buffalo 1.0 is a unified 3D multimodal framework for 3D understanding, text-to-3D generation, instruction-guided 3D editing, and part-level 3D generation. It supports diverse tasks within a single pipeline by connecting language, 3D representations, and generative 3D modules.
The framework uses a shared Hunyuan3D-VLM backbone to bridge 3D QA, grounding, generation, editing, and part-level decomposition. Based on this unified representation, Hunyuan3D DiT modules further enable scalable multimodal generation, high-quality 3D editing, and part generation.
3D Understanding
3D QA and grounding with language-guided reasoning.
Text-to-3D
Generate 3D assets from text prompts.
3D Editing
Edit 3D objects with natural-language instructions.
Part Generation
Extact semantic parts via language.
Compare the source mesh and the edited mesh under user-specified editing instructions.
Case 1 of 1 · Hover a thumbnail to preview its edited result.
Browse generated 3D assets by Hunyuan3D-Buffalo 1.0.
Visualize the whole mesh together with all part components, shown either combined or exploded.