详细信息
- 来源站点
- ArXiv CS.CV
- 作者
- Yanbin Hu, Jin Cui, Jun Ye, Jiepeng Zhou, Jiangcheng Song, Boran Zhao, Pengju Ren
- 文章类型
- PAPER
- 语言
- en
- 发布日期
- 2026-08-04
摘要
arXiv:2608.00110v1 Announce Type: new Abstract: 3D scene understanding requires reasoning about entity existence, spatial layout, and object relations, yet RGB images alone often provide insufficient 3D cues. Existing 3D-VLMs commonly rely on depth or 3D-position-aware inputs at inference time, introducing additional acquisition, reconstruction, or annotation costs that limit RGB-only deployment. We therefore study how training-time 3D evidence can be converted into spatial reasoning capabilities retained under RGB-only inference. We propose a privileged-evidence distillation framework that constructs a distillable teacher through a unified evidence interface and controlled residual injection, and transfers its knowledge to a deployable student receiving only RGB images and questions through logit and structured representation distillation.