详细信息
- 来源站点
- ArXiv CS.CV
- 作者
- Hengyuan Zhang, Jingna Sun, Meiguang Jin, Junfeng Ma
- 文章类型
- PAPER
- 语言
- en
- 发布日期
- 2026-07-28
摘要
arXiv:2607.24013v1 Announce Type: new Abstract: Production-ready audio-driven avatar generation requires efficient inference without sacrificing fidelity or motion expressiveness. However, existing acceleration methods often compromise quality through restrictive architectural choices, such as causal attention and short temporal horizons, or by reducing model capacity and resolution. Without such compromises, we propose AptAvatar, a 14B-parameter long-form audio-driven avatar generation framework that delivers fast and expressive inference. For efficiency in production-level applications, AptAvatar addresses the extreme two-step generation challenge. To bridge the gap between the multi-step teacher model and the two-step student model, we introduce Endpoint-Anchored Distribution Distillation. It augments vanilla distribution matching with a dedicated Anchor Score Estimator trained on the trajectory-endpoint distribution defined from a frozen pretrained 4-step bridge generator.