Pop-Up Distractions Reveal Bag-of-Events Behavior in Video Large Language Models 文章

ArXiv CS.CV2026-08-13PAPERen作者: Oscar Chew, Serhii Honcharenko, Qian-Hui Chen, Patricia Lu, Dishant Zaveri, Khoa D. Doan, Kuan-Hao Huang

详细信息

来源站点
ArXiv CS.CV
作者
Oscar Chew, Serhii Honcharenko, Qian-Hui Chen, Patricia Lu, Dishant Zaveri, Khoa D. Doan, Kuan-Hao Huang
文章类型
PAPER
语言
en
发布日期
2026-08-13

摘要

arXiv:2605.27101v2 Announce Type: replace Abstract: A key capability for video understanding is reliably linking subjects to events across time, yet whether Video Large Language Models (VideoLLMs) actually achieve this remains unclear. In this work, we introduce DistractionBench to evaluate whether VideoLLMs can robustly link subjects and events in the presence of unrelated video segments. Through controlled interventions, such as inserting short advertisement clips into longer videos, we show that VideoLLMs frequently hallucinate interactions between entities from different segments, incorrectly attributing actions from injected advertisements to subjects in the main video. We characterize this systematic hallucination as bag-of-events (BoE) behavior, where models process videos as collections of events rather than temporally structured sequences. Evaluating 11 popular VideoLLMs, we find that all models exhibit substantial BoE behavior.

相关事件

暂无数据

相关公司查看全部 (6)

A
ActuaNONPROFIT
A
ANDINONPROFIT
A
ACTIONNONPROFIT

相关人物

暂无数据