Attention-Guided Saliency Maps for Interpreting Visualization Literacy in VLMs 文章

ArXiv CS.CV2026-07-20PAPERen作者: Maeve Hutchinson, Abderrahmane Wassim Mehdaoui, Pranava Madhyastha

详细信息

来源站点
ArXiv CS.CV
作者
Maeve Hutchinson, Abderrahmane Wassim Mehdaoui, Pranava Madhyastha
文章类型
PAPER
语言
en
发布日期
2026-07-20

摘要

arXiv:2607.16105v1 Announce Type: new Abstract: Understanding how vision-language models (VLMs) interpret data visualizations remains an open problem, and is increasingly important as these models are used for analytical tasks where reliable reasoning is essential. We introduce a lightweight, diagnostic saliency map method tailored for text generation over images using transformer models, the current state-of-the-art models in visualization interpretation. Our approach aggregates the language model's attention over the visual tokens across all heads and layers, then maps this attention back onto the vision encoder's patch grid to localise it over the image, producing a direct correspondence between each generated answer token and the image regions it attended to. This yields fast, gradient-free saliency maps that expose how VLMs allocate focus across visual elements during answer generation, enabling inspection of whether model attention aligns with semantically relevant components.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据

相关产品

暂无数据