From HMM's to segment models: a unified view of stochastic modeling for speech recognition 论文

1996IEEE Transactions on Speech and Audio Processing引用 594
Speech Recognition and SynthesisSpeech and Audio ProcessingMusic and Audio Processing

详细信息

发表期刊/会议
IEEE Transactions on Speech and Audio Processing
发表日期
1996-01-01
发表年份
1996

关键词

Speech Recognition and SynthesisSpeech and Audio ProcessingMusic and Audio Processing

摘要

Many alternative models have been proposed to address some of the shortcomings of the hidden Markov model (HMM), which is currently the most popular approach to speech recognition. In particular, a variety of models that could be broadly classified as segment models have been described for representing a variable-length sequence of observation vectors in speech recognition applications. Since there are many aspects in common between these approaches, including the general recognition and training problems, it is useful to consider them in a unified framework. The paper describes a general stochastic model that encompasses most of the models proposed in the literature, pointing out similarities of the models in terms of correlation and parameter tying assumptions, and drawing analogies between segment models and HMMs. In addition, we summarize experimental results assessing different modeling assumptions and point out remaining open questions.