Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment 文章

ArXiv CS.CV2026-07-07PAPERen作者: Zixiang Zhou, Zhentao Yu, Yifeng Ma, Hongmei Wang, Wenqing Yu, Cong Wang, Zilin Yang, Rui Chen, Jiarong Ou, Yezhou Liu, Yuan Zhou, Qinglin Lu

详细信息

来源站点
ArXiv CS.CV
作者
Zixiang Zhou, Zhentao Yu, Yifeng Ma, Hongmei Wang, Wenqing Yu, Cong Wang, Zilin Yang, Rui Chen, Jiarong Ou, Yezhou Liu, Yuan Zhou, Qinglin Lu
文章类型
PAPER
语言
en
发布日期
2026-07-07

摘要

arXiv:2607.04311v1 Announce Type: new Abstract: Subject-driven and multi-element video generation are central to controllable video synthesis, but existing methods still struggle to preserve identity consistency and model complex relationships among multiple subjects. In this paper, we propose Aura, a unified framework for high-fidelity and identity-consistent video generation. To better capture scene dynamics and subject interactions, we introduce AI director-level captions that provide dense and structured descriptions of video content. We further leverage a vision-language model (VLM) with learnable queries to extract multimodal semantic features from textual and visual references, covering both global semantics and fine-grained visual cues. To bridge the representational gap between the VLM and the Diffusion Transformer (DiT), we design a two-stage alignment strategy that progressively maps VLM features into the DiT feature space.

相关事件

暂无数据

相关公司查看全部 (3)

A
ACTIONNONPROFIT
A
AMI团队RESEARCH_INSTITUTE
A
Advanced Photon SourceRESEARCH_INSTITUTE

相关人物

暂无数据