Calibrate Before Reason: Robust Visual Token Reduction against Semantic Drift in VLMs 文章

ArXiv CS.CV2026-07-31PAPERen作者: Jiasheng Li, Zhong Ji, Yan Zhang, Huihui Li

详细信息

来源站点
ArXiv CS.CV
作者
Jiasheng Li, Zhong Ji, Yan Zhang, Huihui Li
文章类型
PAPER
语言
en
发布日期
2026-07-31

摘要

arXiv:2607.27700v1 Announce Type: new Abstract: Large Vision-Language Models (VLMs) suffer from prohibitive inference overhead due to long sequences of visual tokens. However, existing visual token reduction methods mainly improve efficiency by pruning or compressing redundant tokens without examining whether the resulting representation remains semantically consistent with the original representation. Mapping the original N-token visual sequence to K tokens may discard, dilute, or misassign critical visual cues, triggering severe semantic drift that deviates the VLM's understanding. In this paper, we first introduce the principle of 'Calibrate Before Reason' to visual token reduction and propose CaRe, a training-free robust framework that calibrates compact visual representations before reasoning to preserve semantic fidelity in VLMs.

相关事件

暂无数据

相关公司查看全部 (2)

A
ANDINONPROFIT
A
AMI团队RESEARCH_INSTITUTE

相关人物

暂无数据