StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring 文章

ArXiv CS.CV2026-08-04PAPERen作者: Xiangyue Zhang, Jianfang Li, Jiaxu Zhang, Kaixing Yang, Steven Hoi

详细信息

来源站点
ArXiv CS.CV
作者
Xiangyue Zhang, Jianfang Li, Jiaxu Zhang, Kaixing Yang, Steven Hoi
文章类型
PAPER
语言
en
发布日期
2026-08-04

摘要

arXiv:2608.01643v1 Announce Type: new Abstract: Real-time co-speech gesture generation must produce 3D motion clip by clip as speech arrives. Existing streaming methods are open-loop: each clip depends on past context, but the model cannot check or correct its trajectory. Small errors therefore accumulate and cause drift over long sequences. We observe that this failure is mainly caused by the lack of a forward constraint rather than poor short-clip quality. A plausible key pose at the end of each clip provides a destination anchor that limits drift. Based on this observation, we propose StreamTalk, a closed-loop framework with a periodic generate-retrieve-refine cycle. Streaming Pose-Guided Generation first predicts a coarse clip, retrieves a plausible tail pose from a speaker-specific motion database, and refines the clip using this pose before continuing to the next window.

相关事件

暂无数据

相关公司查看全部 (4)

A
ATHCOMPANY
A
AT TCOMPANY
A
AMI团队RESEARCH_INSTITUTE

相关人物

暂无数据