DocumentCode :
1467126
Title :
Voice Conversion Using Partial Least Squares Regression
Author :
Helander, Elina ; Virtanen, Tuomas ; Nurminen, Jani ; Gabbouj, Moncef
Author_Institution :
Dept. of Signal Process., Tampere Univ. of Technol., Tampere, Finland
Volume :
18
Issue :
5
fYear :
2010
fDate :
7/1/2010 12:00:00 AM
Firstpage :
912
Lastpage :
921
Abstract :
Voice conversion can be formulated as finding a mapping function which transforms the features of the source speaker to those of the target speaker. Gaussian mixture model (GMM)-based conversion is commonly used, but it is subject to overfitting. In this paper, we propose to use partial least squares (PLS)-based transforms in voice conversion. To prevent overfitting, the degrees of freedom in the mapping can be controlled by choosing a suitable number of components. We propose a technique to combine PLS with GMMs, enabling the use of multiple local linear mappings. To further improve the perceptual quality of the mapping where rapid transitions between GMM components produce audible artefacts, we propose to low-pass filter the component posterior probabilities. The conducted experiments show that the proposed technique results in better subjective and objective quality than the baseline joint density GMM approach. In speech quality conversion preference tests, the proposed method achieved 67% preference score against the smoothed joint density GMM method and 84% preference score against the unsmoothed joint density GMM method. In objective tests the proposed method produced a lower Mel-cepstral distortion than the reference methods.
Keywords :
Gaussian processes; least squares approximations; low-pass filters; regression analysis; speech processing; Gaussian mixture model-based conversion; Mel-cepstral distortion; component posterior probabilities; low-pass filter; mapping function; multiple local linear mappings; partial least squares regression; reference methods; speech quality conversion preference tests; voice conversion; Gaussian mixture model (GMM); partial least squares regression; voice conversion;
fLanguage :
English
Journal_Title :
Audio, Speech, and Language Processing, IEEE Transactions on
Publisher :
ieee
ISSN :
1558-7916
Type :
jour
DOI :
10.1109/TASL.2010.2041699
Filename :
5445056
Link To Document :
بازگشت