Why We Need Speech to Evaluate Speech Translation 文章

ArXiv CS.CL2026-05-28NEWSen作者: Maike Z\"ufle, Danni Liu, Vil\'em Zouhar, Jan Niehues

详细信息

来源站点: ArXiv CS.CL
作者: Maike Z\"ufle, Danni Liu, Vil\'em Zouhar, Jan Niehues
文章类型: NEWS
语言: en
发布日期: 2026-05-28

摘要

arXiv:2605.28227v1 Announce Type: new Abstract: Speech translation models are increasingly capable of preserving speech-specific information (e.g., speaker gender, prosody, and emphasis), yet evaluation metrics remain blind to such phenomena. We meta-evaluate both text- and speech-based quality estimation metrics on two contrastive datasets targeting gender agreement and prosody, and find that both fall short, even when given direct access to the speech signal. We then train SpeechCOMET, a family of quality estimation models with speech encoders, and evaluate a state-of-the-art SpeechLLM as a judge. Both match or exceed text-based COMET on standard quality estimation, but neither consistently assesses speech-specific phenomena. We identify three causes: (1) speech-specific features are not reliably preserved in current encoders, (2) models tend to ignore the speech source signal, and (3) quality estimation training data contains too few relevant examples.

Why We Need Speech to Evaluate Speech Translation 文章

详细信息

摘要

相关事件

相关公司

相关人物

相关产品查看全部 (3)

相关技术查看全部 (1)