Title :
Unit selection speech synthesis using multiple speech units at non-adjacent segments for prosody and waveform generation
Author :
Tamura, Masatsune ; Braunschweiler, Norbert ; Kagoshima, Takehiko ; Akamine, Masami
Author_Institution :
Corp. R&D Center, Toshiba Corp., Kawasaki, Japan
Abstract :
In this paper, we propose a speech synthesis method that combines a natural waveform concatenation based speech synthesis method and our baseline plural unit selection and fusion method. Two main features of the proposed method are (i) prosody regeneration from selected speech units and (ii) using multiple speech units at non-adjacent segments. The nonadjacent segments is the segment that the previous or following speech units in the optimum speech unit sequence are not adjacent in the database. By using the prosody of selected speech units, the original prosodic expressions and sounds of recorded speech are retained, while discontinuities are reduced by using multiple speech units at non-adjacent segments. MOS evaluations showed that the proposed method provides a clear improvement against the conventional unit selection method and our baseline method.
Keywords :
linguistics; speech synthesis; waveform analysis; MOS evaluation; baseline plural unit selection; fusion method; natural waveform concatenation; non adjacent segments; prosody regeneration; speech synthesis method; unit selection; Cost function; Degradation; Europe; Frequency; Fusion power generation; Laboratories; Natural languages; Research and development; Spatial databases; Speech synthesis; concatenative speech synthesis; prosody generation; unit fusion; unit selection;
Conference_Titel :
Acoustics Speech and Signal Processing (ICASSP), 2010 IEEE International Conference on
Conference_Location :
Dallas, TX
Print_ISBN :
978-1-4244-4295-9
Electronic_ISBN :
1520-6149
DOI :
10.1109/ICASSP.2010.5495151