• DocumentCode
    2020248
  • Title

    Tree-based unit selection for English speech synthesis

  • Author

    Wang, Wern Jun ; Campbell, W.N. ; Iwahashi, Naoto ; Sagisaka, Yoshinori

  • Author_Institution
    ATR Interpreting Telephony Res. Lab., Kyoto, Japan
  • Volume
    2
  • fYear
    1993
  • fDate
    27-30 April 1993
  • Firstpage
    191
  • Abstract
    In concatenative synthesis for English, scarcity of speech data for many contexts is a serious problem. The authors propose a novel unit selection scheme using a decision-tree-based clustering method that combines acoustic and linguistic knowledge with statistical modeling. This approach makes it possible not only to find a trainable and consistent set of generalized allophonic models but also to achieve some local optimality with respect to the limited training data. To evaluate the validity of this algorithm, regression tree generation has been carried out for both vowels and consonants from 200 phonetically balanced sentences read by a female speaker. Experimental results show that regression trees offer a promising solution for the data scarcity problem.<>
  • Keywords
    audio acoustics; decision theory; knowledge based systems; speech synthesis; trees (mathematics); English speech synthesis; acoustic knowledge; algorithm; data scarcity; decision-tree-based clustering method; generalized allophonic models; linguistic knowledge; regression tree generation; statistical modeling; training data; unit selection scheme;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Acoustics, Speech, and Signal Processing, 1993. ICASSP-93., 1993 IEEE International Conference on
  • Conference_Location
    Minneapolis, MN, USA
  • ISSN
    1520-6149
  • Print_ISBN
    0-7803-7402-9
  • Type

    conf

  • DOI
    10.1109/ICASSP.1993.319266
  • Filename
    319266