• DocumentCode
    3430524
  • Title

    Multi-speaker prosodic instance selection for HMM-based speech synthesis

  • Author

    Yansuo Yu ; Fengyun Zhu ; Xihong Wu

  • Author_Institution
    Key Lab. of Machine Perception (Minist. of Educ.), Peking Univ., Beijing, China
  • fYear
    2013
  • fDate
    6-10 July 2013
  • Firstpage
    142
  • Lastpage
    146
  • Abstract
    In this paper, we propose a novel parametric speech synthesis based on prosodic instance selection to improve the naturalness of synthesized speech especially in the case of small database. Prosodic instances including F0 and duration are directly selected from the database to preserve rich prosodic variations other than generation from the statistical models. Considering that spectral and prosodic parameters could be modeled separately, prosodic instances from multiple speakers, which are easier to obtain than those of single speaker, are exploited to not only enhance the prosodic models but also enrich the coverage of prosodic context for the synthesized speaker. The results of subjective listening tests on the corresponding databases further show that the proposed method can achieve better performance than both parametric synthesis and waveform concatenation synthesis.
  • Keywords
    hidden Markov models; speech synthesis; HMM-based speech synthesis; hidden Markov model; multispeaker prosodic instance selection; novel parametric speech synthesis; parametric synthesis; prosodic variations; statistical models; subjective listening tests; waveform concatenation synthesis; Context modeling; Databases; Hidden Markov models; Mathematical model; Speech; Speech synthesis; Training; HMM-Based Speech Synthesis; Prosodic Instance Selection (PIS);
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Signal and Information Processing (ChinaSIP), 2013 IEEE China Summit & International Conference on
  • Conference_Location
    Beijing
  • Type

    conf

  • DOI
    10.1109/ChinaSIP.2013.6625315
  • Filename
    6625315