DDVT: Dynamic Dual-level Vision Transformer Fusion Network for Answer Grounding in Visual Question Answering 文章

ArXiv CS.CV2026-07-28PAPERen作者: Yue Zhang, Xiangyu Li, Wanshu Fan, Xin Yang, Dongsheng Zhou

详细信息

来源站点
ArXiv CS.CV
作者
Yue Zhang, Xiangyu Li, Wanshu Fan, Xin Yang, Dongsheng Zhou
文章类型
PAPER
语言
en
发布日期
2026-07-28

摘要

arXiv:2607.23921v1 Announce Type: new Abstract: Answer grounding in visual question answering aims to locate the region from a given natural language question associated with the visual content of an image, which has garnered significant attention due to its practical applications. In this paper, we introduce the Dynamic Dual-level Vision Transformer Fusion Network (DDVT) for answer grounding in visual question answering. Specifically, we propose a question-guided dynamic regional-level module (QGDR) that combines complementary image context through ROI Align and text content, enabling precise localization of text-related visual content. Moreover, we present a cross-modal multi-scale aggregation module (CMA) that enhances feature fusion between pixel-level and region-level features, facilitating the effective localization of visual content associated with grounded answers.