Disentangling Semantic Attention from Structural Bias in the Attention Manifold 文章

ArXiv CS.CV2026-07-28PAPERen作者: Pengkun Jiao, Bin Zhu, Jingjing Chen, Yu-gang Jiang

详细信息

来源站点
ArXiv CS.CV
作者
Pengkun Jiao, Bin Zhu, Jingjing Chen, Yu-gang Jiang
文章类型
PAPER
语言
en
发布日期
2026-07-28

摘要

arXiv:2607.24017v1 Announce Type: new Abstract: The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws. Specifically, MLLMs consistently exhibit disproportionate attention toward certain semantically uninformative visual tokens, a phenomenon termed "register" or "Visual Attention Sinks." While existing inference intervention methods attempt to identify these sink tokens and redistribute their attention weights, such approaches typically treat these tokens in isolation and suffer from computational inefficiency. Instead, we reframe this phenomenon as a generalized textual bias exerted over visual features that extends beyond isolated sink tokens. From this perspective, a pervasive structural bias leads to the dilution of the semantic visual signal, precipitating multimodal hallucinations as the model prioritizes linguistic priors over valid visual evidence.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据

相关产品

暂无数据