Streaming Neural Speech Codecs through Time-Invariant Representations 文章

ArXiv CS.CL2026-07-07PAPERen作者: K\'elian Est\`eve, Salima Mhdaffar, Mickael Rouvier, Richard Dufour, Yannick Est\`eve

详细信息

来源站点
ArXiv CS.CL
作者
K\'elian Est\`eve, Salima Mhdaffar, Mickael Rouvier, Richard Dufour, Yannick Est\`eve
文章类型
PAPER
语言
en
发布日期
2026-07-07

摘要

arXiv:2607.05250v1 Announce Type: new Abstract: Neural speech codecs are increasingly used as intermediate representations in codec-based speech generation systems. TiCodec introduces a factorized representation that separates time-varying speech content from time-invariant information through a Time-Invariant Representation Extraction (TIRE) module, potentially reducing the amount of information that must be modeled at the frame-level. In this work, we investigate the nature of the information captured by TIRE representations and their suitability for low-latency speech processing. Using a series of probing tasks, we analyze the influence of the encoder layer and show that intermediate layers capture complementary speaker- and environment-related information while containing little linguistic content. We further study several segment selection strategies for TIRE training and demonstrate that cross-file sampling improves the robustness of invariant representations.

相关事件

暂无数据

相关公司查看全部 (5)

A
AT TCOMPANY
A
AMI团队RESEARCH_INSTITUTE

相关人物

暂无数据