详细信息
- 来源站点
- ArXiv CS.CV
- 作者
- Gengtian Shi, Jinze Yu, Chenhao Wu, Shaofei Wang, Eiji Fukuzawa, Junjie Tang, Hiroshi Onoda, Jiang Liu
- 文章类型
- PAPER
- 语言
- en
- 发布日期
- 2026-07-07
别名
摘要
arXiv:2607.05093v1 Announce Type: new Abstract: The exponential growth of video traffic demands novel semantic communication paradigms that transmit meaning rather than raw bits. We present a generative AI-enabled framework for semantic video communication addressing two critical challenges: efficient hierarchical temporal modeling for bandwidth-constrained transmission and robust semantic alignment between video content and natural language queries at network edge devices. Our approach introduces a multi-scale temporal convolutional encoder that captures motion patterns across different temporal granularities with O(T) complexity suitable for resource-constrained IoT deployments. We further propose a capsule-based dynamic routing mechanism that iteratively refines segment-query associations, enabling flexible modeling of non-monotonic semantic alignments essential for goal-oriented communication.