DocumentCode
846128
Title
Bayesian learning of speech duration models
Author
Chien, Jen-Tzung ; Huang, Chih-Hsien
Author_Institution
Dept. of Comput. Sci. & Inf. Eng., Nat. Cheng Kung Univ., Tainan, Taiwan
Volume
11
Issue
6
fYear
2003
Firstpage
558
Lastpage
567
Abstract
This paper presents the Bayesian speech duration modeling and learning for hidden Markov model (HMM) based speech recognition. We focus on the sequential learning of HMM state duration using quasi-Bayes (QB) estimate. The adapted duration models are robust to nonstationary speaking rates and noise conditions. In this study, the Gaussian, Poisson, and gamma distributions are investigated to characterize the duration models. The maximum a posteriori (MAP) estimate of gamma duration model is developed. To exploit the sequential learning, we adopt the Poisson duration model incorporated with gamma prior density, which belongs to the conjugate prior family. When the adaptation data are sequentially observed, the gamma posterior density is produced with twofold advantages. One is to determine the optimal QB duration parameter, which can be merged in HMMs for speech recognition. The other one is to build the updating mechanism of gamma prior statistics for sequential learning. EM algorithm is applied to fulfill QB parameter estimation. The adaptation of overall HMM parameters can be performed simultaneously. In the experiments, the proposed adaptive duration model improves the speech recognition performance of Mandarin broadcast news and noisy connected digits. The batch and sequential learning are respectively investigated for MAP and QB duration models.
Keywords
Bayes methods; Gaussian distribution; Poisson distribution; gamma distribution; hidden Markov models; learning systems; maximum likelihood estimation; speech recognition; Bayesian learning; Gaussian distribution; Mandarin broadcast news; Poisson distribution; adaptive duration model; gamma distribution; gamma posterior density; hidden Markov model; maximum a posteriori estimation; noisy connected digits; parameter estimation; quasiBayes estimation; sequential learning; speaking rate; speech duration model; speech recognition; state duration; Automatic speech recognition; Bayesian methods; Broadcasting; Decoding; Hidden Markov models; Humans; Speech recognition; Speech synthesis; Statistics; Vocabulary;
fLanguage
English
Journal_Title
Speech and Audio Processing, IEEE Transactions on
Publisher
ieee
ISSN
1063-6676
Type
jour
DOI
10.1109/TSA.2003.818114
Filename
1255444
Link To Document