详细信息
- 来源站点
- ArXiv CS.CV
- 作者
- Yuwen Tan, Joey Huang, Jin Huang, Haoxiang Li, Boqing Gong
- 文章类型
- PAPER
- 语言
- en
- 发布日期
- 2026-07-07
摘要
arXiv:2607.03043v1 Announce Type: new Abstract: Understanding camera movement in natural language is critical for training and evaluating video generation models, among other applications. However, we demonstrate that existing vision-language models (VLMs) fail this task in surprising ways, frequently confusing translation with rotation, left with right, and object movement with camera movement. To address these limitations, we establish natural language camera movement understanding as a standalone research task. We introduce a two-level cinematographic taxonomy and an extensive, atomic benchmark featuring both real and synthetic videos. Furthermore, we curate a large-scale, multi-source training set enhanced by targeted camera movement augmentation. Our fine-tuned VLM-8B outperforms Gemini 3.1 Pro by 10% and 11% on our benchmark's real and synthetic videos, respectively.
相关事件
暂无数据
相关人物
暂无数据