UMA-Split: unimodal aggregation for both English and Mandarin non-autoregressive speech recognition 文章

ArXiv CS.CL2026-06-18NEWSen作者: Ying Fang, Xiaofei Li

详细信息

来源站点
ArXiv CS.CL
作者
Ying Fang, Xiaofei Li
文章类型
NEWS
语言
en
发布日期
2026-06-18

摘要

arXiv:2509.14653v2 Announce Type: replace Abstract: This paper proposes a unimodal aggregation (UMA) based nonautoregressive model for both English and Mandarin speech recognition. The original UMA explicitly segments and aggregates acoustic frames (with unimodal weights that first monotonically increase and then decrease) of the same text token to learn better representations than regular connectionist temporal classification (CTC). However, it only works well in Mandarin. It struggles with other languages, such as English, for which a single syllable may be tokenized into multiple fine-grained tokens, or a token spans fewer than 3 acoustic frames and fails to form unimodal weights. To address this problem, we propose allowing each UMA-aggregated frame map to multiple tokens, via a simple split module that generates two tokens from each aggregated frame before computing the CTC loss.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据

相关产品

暂无数据