Inference-Time Agentic Decision Rules Beat Longer Evolving Search for Multi-Image Medical Reasoning 文章

ArXiv CS.CV2026-07-31PAPERen作者: Site Li, Jianyi Hao, Xiaofeng Liu

详细信息

来源站点
ArXiv CS.CV
作者
Site Li, Jianyi Hao, Xiaofeng Liu
文章类型
PAPER
语言
en
发布日期
2026-07-31

摘要

arXiv:2607.27564v1 Announce Type: new Abstract: Multi-image medical VQA is not merely a prompt-length problem; it is a fundamental challenge of agentic decision-making. Medical vision-language agents must aggregate evidence across ordered images, remain robust to answer-order perturbations, and avoid overfitting to noisy search-time feedback. We study MedFrameQA through a controlled comparison of five inference-time agentic strategies, optimized using the same high-budget ShinkaEvolve configuration and evaluated on a reproducible internal frozen split (1,331 evolution, 665 holdout, 855 final test). Across five independent repeated runs, the strongest method emerges as the simplest robust aggregator: the \textbf{order-vote} policy achieves $57.89 \pm 0.65\%$ final-test accuracy, significantly outperforming the fixed baseline ($52.73 \pm 0.42\%$) and the more complex, albeit brittle, order-rerank variant ($55.79 \pm 0.43\%$). Paired bootstrap analysis confirms these significant gains.