Trace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs 文章

ArXiv CS.CV2026-08-06PAPERen作者: Yang Yang, Jiawei Chen, Tairan Chen, Zhaoxia Yin

详细信息

来源站点
ArXiv CS.CV
作者
Yang Yang, Jiawei Chen, Tairan Chen, Zhaoxia Yin
文章类型
PAPER
语言
en
发布日期
2026-08-06

摘要

arXiv:2608.04759v1 Announce Type: new Abstract: Although Multimodal Large Language Models (MLLMs) have made substantial progress, their spatial reasoning may still produce intermediate judgments inconsistent with the input image, allowing errors to propagate through the reasoning chain and affect the final answer. Existing methods mainly improve spatial reasoning through training or additional spatial information, without considering whether the reasoning process itself is faithful to the model input. Our study shows that unfaithful reasoning chains significantly reduce final-answer accuracy. To address this issue, we propose a modular and training-free framework for spatial reasoning verification and correction. The framework constructs a Spatial Evidence Graph (SEG), which associates atomic spatial evidence extracted from Chain-of-Thought reasoning with visual entities, spatial relations, source steps, and visual evidence.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据