MiniWorld: Democratizing the Training of Video World Models from Scratch 文章

ArXiv CS.CV2026-08-04PAPERen作者: Yian Zhao, Ruochong Zheng, Hongcan Guo, Yu Yan, Jian Zhang, Jie Chen

详细信息

来源站点
ArXiv CS.CV
作者
Yian Zhao, Ruochong Zheng, Hongcan Guo, Yu Yan, Jian Zhang, Jie Chen
文章类型
PAPER
语言
en
发布日期
2026-08-04

摘要

arXiv:2608.01127v1 Announce Type: new Abstract: Video world models predict future observations conditioned on historical observations and control signals, enabling long-horizon generation through autoregressive state transitions. Unlike conventional video generation models that primarily capture visual appearance and motion, video world models learn the underlying dynamics governing environment evolution under agent actions, providing a foundation for embodied AI and interactive simulation. Recent progress has largely relied on adapting pretrained video generation models through post-training or distillation. Although effective, these approaches often require complex training pipelines, substantial computational resources, and suffer from the mismatch between bidirectional pretraining and causal streaming inference. Recent studies have shown that training autoregressive video world models from scratch is feasible and scalable.

相关事件

暂无数据

相关公司查看全部 (4)

A
AT TCOMPANY
A
AMI团队RESEARCH_INSTITUTE
A
ACTIONNONPROFIT

相关人物

暂无数据