FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification 文章

ArXiv CS.CV2026-07-31PAPERen作者: Haoqing Wang, Xingrun Xing, Wei Xia, Ziheng Li, Yehui Tang

详细信息

来源站点
ArXiv CS.CV
作者
Haoqing Wang, Xingrun Xing, Wei Xia, Ziheng Li, Yehui Tang
文章类型
PAPER
语言
en
发布日期
2026-07-31

摘要

arXiv:2607.28225v1 Announce Type: new Abstract: Agentic vision-language models (VLMs), which interleave textual reasoning with explicit tool calls such as cropping and code-based image manipulation, have emerged as a compelling paradigm for reliable and interpretable multimodal reasoning. However, recent studies have revealed that such models often use tools unfaithfully. Many process images are irrelevant to the question (e.g., the tool crops the wrong region or misses the queried target), yet the call still receives full credit and the model still answers correctly. Such decorative or misaligned tool calls waste computation and reveal that the model leans on prior knowledge or the original image rather than the evidence it retrieves. This may stem from two limitations of prevailing methods: the tool reward fails to distinguish useful from useless calls, and tool feedback carries no signal of usefulness. To this end, we introduce FaithEyes, a multi-agent self-judging framework.

相关事件

暂无数据

相关公司查看全部 (5)

A
ANICOMPANY
A
AT TCOMPANY
A
ATHCOMPANY

相关人物

暂无数据