Natural Language Camera Movement Understanding 文章

ArXiv CS.CV2026-07-07PAPERen作者: Yuwen Tan, Joey Huang, Jin Huang, Haoxiang Li, Boqing Gong

详细信息

来源站点
ArXiv CS.CV
作者
Yuwen Tan, Joey Huang, Jin Huang, Haoxiang Li, Boqing Gong
文章类型
PAPER
语言
en
发布日期
2026-07-07

摘要

arXiv:2607.03043v1 Announce Type: new Abstract: Understanding camera movement in natural language is critical for training and evaluating video generation models, among other applications. However, we demonstrate that existing vision-language models (VLMs) fail this task in surprising ways, frequently confusing translation with rotation, left with right, and object movement with camera movement. To address these limitations, we establish natural language camera movement understanding as a standalone research task. We introduce a two-level cinematographic taxonomy and an extensive, atomic benchmark featuring both real and synthetic videos. Furthermore, we curate a large-scale, multi-source training set enhanced by targeted camera movement augmentation. Our fine-tuned VLM-8B outperforms Gemini 3.1 Pro by 10% and 11% on our benchmark's real and synthetic videos, respectively.