ArtNet: A JEPA-Like Articulatory Predictive Framework for Robust Zero-Shot Phoneme Recognition 文章

ArXiv CS.AI2026-06-16NEWSen作者: Zeqian Hu, Fuliang Weng, Shu Shang, Yaqian Zhou

详细信息

来源站点
ArXiv CS.AI
作者
Zeqian Hu, Fuliang Weng, Shu Shang, Yaqian Zhou
文章类型
NEWS
语言
en
发布日期
2026-06-16

摘要

arXiv:2606.16595v1 Announce Type: cross Abstract: Zero-shot cross-lingual phoneme recognition is often hindered by the fragility of direct acoustic-to-symbol mapping, which is susceptible to language-specific variations. Echoing joint-embedding predictive architecture (JEPA) work in vision, we propose ArtNet, a framework that explores a structured feature prediction task based on articulatory features to enhance acoustic robustness. Specifically, ArtNet integrates an articulatory predictor, designed to extract universal articulatory representations from self-supervised learning (SSL) features, with a variational information bottleneck (VIB) to suppress language-specific variations. Experiments on seven unseen languages demonstrate that ArtNet, particularly when synergized with the proposed vector-space inventory alignment (VSIA) strategy, significantly outperforms competitive baselines, achieving a 20.56\% relative reduction in phoneme error rate (PER) and 7.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据