DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning 文章

ArXiv CS.CV2026-07-28PAPERen作者: Mengqi Zhang, Sahil Khose, Simar Kareer, Yuchen Song, Unnat Jain, Judy Hoffman

详细信息

来源站点
ArXiv CS.CV
作者
Mengqi Zhang, Sahil Khose, Simar Kareer, Yuchen Song, Unnat Jain, Judy Hoffman
文章类型
PAPER
语言
en
发布日期
2026-07-28

摘要

arXiv:2607.24159v1 Announce Type: cross Abstract: Generalizable robot manipulation requires policies that can anticipate how visual scenes evolve while executing language instructions. While recent Vision-Language-Action models benefit from large-scale pretraining, their predominantly static pretraining objectives provide limited supervision for physical dynamics and temporal causality, leaving control-relevant knowledge to be learned from downstream robot demonstrations. Video generative models offer a promising foundation by encoding rich spatiotemporal priors through future predictions. However, existing Video-Action Models either couple video and action prediction in a shared backbone, making policy adaptation harder to optimize, or under-utilize video information when guiding the action branch. In this work, we introduce DeVA, a Decoupled Video-Action model with specialized video and action experts, multi-level feature transfer, and physically salient guidance.

相关事件

暂无数据

相关公司查看全部 (4)

A
ANICOMPANY
A
AMI团队RESEARCH_INSTITUTE
A
ACTIONNONPROFIT

相关人物

暂无数据