• DocumentCode
    846128
  • Title

    Bayesian learning of speech duration models

  • Author

    Chien, Jen-Tzung ; Huang, Chih-Hsien

  • Author_Institution
    Dept. of Comput. Sci. & Inf. Eng., Nat. Cheng Kung Univ., Tainan, Taiwan
  • Volume
    11
  • Issue
    6
  • fYear
    2003
  • Firstpage
    558
  • Lastpage
    567
  • Abstract
    This paper presents the Bayesian speech duration modeling and learning for hidden Markov model (HMM) based speech recognition. We focus on the sequential learning of HMM state duration using quasi-Bayes (QB) estimate. The adapted duration models are robust to nonstationary speaking rates and noise conditions. In this study, the Gaussian, Poisson, and gamma distributions are investigated to characterize the duration models. The maximum a posteriori (MAP) estimate of gamma duration model is developed. To exploit the sequential learning, we adopt the Poisson duration model incorporated with gamma prior density, which belongs to the conjugate prior family. When the adaptation data are sequentially observed, the gamma posterior density is produced with twofold advantages. One is to determine the optimal QB duration parameter, which can be merged in HMMs for speech recognition. The other one is to build the updating mechanism of gamma prior statistics for sequential learning. EM algorithm is applied to fulfill QB parameter estimation. The adaptation of overall HMM parameters can be performed simultaneously. In the experiments, the proposed adaptive duration model improves the speech recognition performance of Mandarin broadcast news and noisy connected digits. The batch and sequential learning are respectively investigated for MAP and QB duration models.
  • Keywords
    Bayes methods; Gaussian distribution; Poisson distribution; gamma distribution; hidden Markov models; learning systems; maximum likelihood estimation; speech recognition; Bayesian learning; Gaussian distribution; Mandarin broadcast news; Poisson distribution; adaptive duration model; gamma distribution; gamma posterior density; hidden Markov model; maximum a posteriori estimation; noisy connected digits; parameter estimation; quasiBayes estimation; sequential learning; speaking rate; speech duration model; speech recognition; state duration; Automatic speech recognition; Bayesian methods; Broadcasting; Decoding; Hidden Markov models; Humans; Speech recognition; Speech synthesis; Statistics; Vocabulary;
  • fLanguage
    English
  • Journal_Title
    Speech and Audio Processing, IEEE Transactions on
  • Publisher
    ieee
  • ISSN
    1063-6676
  • Type

    jour

  • DOI
    10.1109/TSA.2003.818114
  • Filename
    1255444