DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents 文章

ArXiv CS.CV2026-08-04PAPERen作者: Tongsheng Ding, Zhen Luo, Yixuan Yang, Boyu Wang, Luyang Xie, Jinyu Yang, Feng Zheng

详细信息

来源站点
ArXiv CS.CV
作者
Tongsheng Ding, Zhen Luo, Yixuan Yang, Boyu Wang, Luyang Xie, Jinyu Yang, Feng Zheng
文章类型
PAPER
语言
en
发布日期
2026-08-04

摘要

arXiv:2608.00486v1 Announce Type: new Abstract: Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. Progress is limited on two fronts: available datasets lack fine-grained language-to-motion annotations, and existing predictors either rely on privileged inputs such as video, depth, or CAD models, or recover motion from fully generated videos through costly, error-prone perception pipelines. We close the supervision gap with the MOVE dataset, 5,038 object-centric egocentric trajectories, each paired with a fine-grained natural-language instruction rather than a coarse verb-noun label. We further propose DreamTraj, which predicts a 6-DoF object trajectory from a single RGB image and a task instruction, requiring no video, depth, or CAD model at inference: rather than generating a video, it reads motion from the internal representations of a frozen image-to-video diffusion model at an early denoising step.

相关事件

暂无数据

相关公司查看全部 (5)

A
ATHCOMPANY
A
ANICOMPANY

相关人物

暂无数据