MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models 文章

ArXiv CS.CV2026-07-31PAPERen作者: Wenjie Zhu, Yabin Zhang, Wenjun Zeng, Lei Zhang

详细信息

来源站点
ArXiv CS.CV
作者
Wenjie Zhu, Yabin Zhang, Wenjun Zeng, Lei Zhang
文章类型
PAPER
语言
en
发布日期
2026-07-31

摘要

arXiv:2607.27637v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or shifted contexts. A reliable MLLM should refuse truly out-of-context (OOC) questions with subject-level context shifts while still answering shifted in-context (Shifted IC) questions with non-subject context shifts. Existing benchmarks mainly target OOC or visually unanswerable questions, but overlook answerable Shifted IC cases and cover limited OOC shifts. To fill this gap, we present MMOOC, a large-scale benchmark for evaluating refusal and robust answering abilities of MLLMs. MMOOC contains over 41K image-question pairs, including answerable Shifted IC cases and unanswerable OOC cases, spanning three question formats, eight shift types and six visual scenarios, with data quality ensured through MLLM-based filtering and human verification.