GuideGround: VLM-guided Semantic Understanding and Viewpoint-aware Reasoning for 3D Visual Grounding 文章

ArXiv CS.CV2026-08-04PAPERen作者: Yiwen Wang, Yuyang Deng, Yihao Long, Xi Zhao

详细信息

来源站点
ArXiv CS.CV
作者
Yiwen Wang, Yuyang Deng, Yihao Long, Xi Zhao
文章类型
PAPER
语言
en
发布日期
2026-08-04

摘要

arXiv:2608.00518v1 Announce Type: new Abstract: 3D visual grounding aims to localize the target object in a 3D scene from a natural language query, requiring both fine-grained semantic understanding and viewpoint-dependent spatial reasoning. Existing methods typically formulate semantic understanding as an auxiliary closed-set object classification task and rely on multi-view feature aggregation for viewpoint reasoning, limiting semantic generalization and weakening viewpoint-specific evidence. We observe that vision-language models naturally provide complementary capabilities through open-vocabulary semantic understanding and global scene perception. Based on this insight, we propose GuideGround, a VLM-guided framework that complements rather than replaces task-specific grounding models by leveraging VLMs for semantic enhancement and viewpoint-specific hypothesis verification.