Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin 文章

ArXiv CS.CL2026-07-24PAPERen作者: Zhiheng Qian, Aini Li, Hai Hu, Liang Zhao

详细信息

来源站点
ArXiv CS.CL
作者
Zhiheng Qian, Aini Li, Hai Hu, Liang Zhao
文章类型
PAPER
语言
en
发布日期
2026-07-24

摘要

arXiv:2607.21332v1 Announce Type: new Abstract: Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties. We address this by training text-dependent and text-independent aligners for Chengdu Mandarin using a 17-hour corpus and a custom G2P dictionary. We trained a text-dependent GMM-HMM model (Chengdu-MFA) and fine-tuned a pretrained audio encoder on frame classification with Chengdu-MFA's pseudo label for text-independent alignment (Chengdu-FC). Evaluation on an expert-annotated test set show that both methods significantly outperform Standard Mandarin baselines. Chengdu-MFA reduced average phone boundary differences by 31.8%, while Chengdu-FC achieved a 61.2% reduction. This work establishes a practical bootstrapping pipeline for developing accurate aligners for under-resourced varieties without labor- and time-intensive manual annotation.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据