Holo-Captioning: Toward the Text Equivalent of 3D Scenes 文章

ArXiv CS.CV2026-07-07PAPERen作者: Kun-Yu Lin, Chengke Bu, Zhenguo Li, Kai Han

详细信息

来源站点
ArXiv CS.CV
作者
Kun-Yu Lin, Chengke Bu, Zhenguo Li, Kai Han
文章类型
PAPER
语言
en
发布日期
2026-07-07

摘要

arXiv:2607.02908v1 Announce Type: new Abstract: This work introduces holo-captioning, a novel task that strives to seek the text equivalent of 3D scenes. As the initial step, we formulate holo-captioning as generating a structured textual description that comprehensively depicts all entities within a 3D scene -- including their semantic tags, spatial locations, attributes, and inter-entity relations. To tackle this challenging task, we first develop an effective captioning engine to produce detailed descriptions of individual entity instances and instance pairs, and contribute a large-scale benchmark comprising over 15K scenes for training and evaluation. Building upon this foundation, we propose HoloScribe, a novel model that features an instance-aware decoupled pipeline for generating structured holo-captions, and further incorporates anchor-aware instance linking to identify relational instance pairs.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据