Enriching speech recognition with automatic detection of sentence boundaries and disfluencies 论文

2006IEEE Transactions on Audio Speech and Language Processing引用 270

Natural Language Processing TechniquesSpeech Recognition and SynthesisSpeech and dialogue systems

Natural Language Processing Techniques Speech Recognition and Synthesis Speech and dialogue systems

作者

摘要

Effective human and automatic processing of speech requires recovery of more than just the words. It also involves recovering phenomena such as sentence boundaries, filler words, and disfluencies, referred to as structural metadata. We describe a metadata detection system that combines information from different types of textual knowledge sources with information from a prosodic classifier. We investigate maximum entropy and conditional random field models, as well as the predominant hidden Markov model (HMM) approach, and find that discriminative models generally outperform generative models. We report system performance on both broadcast news and conversational telephone speech tasks, illustrating significant performance differences across tasks and as a function of recognizer performance. The results represent the state of the art, as assessed in the NIST RT-04F evaluation

作者查看全部 (5)

Mary P. Harper

Mari Ostendorf

Dustin Hillard

Andreas Stolcke

Enriching speech recognition with automatic detection of sentence boundaries and disfluencies 论文

摘要

作者查看全部 (5)

相关技术查看全部 (2)

相关事件

相关文章