DocumentCode
3716067
Title
Evaluation of PNCC and extended spectral subtraction methods for robust speech recognition
Author
Thibaut Fux;Denis Jouvet
Author_Institution
Inria, 615 rue du Jardin Botanique, F-54600, Villers-lè
fYear
2015
Firstpage
1416
Lastpage
1420
Abstract
This paper evaluates the robustness of different approaches for speech recognition with respect to signal-to-noise ratio (SNR), to signal level and to presence of non-speech data before and after utterances to be recognized. Three types of noise robust features are considered: Power Normalized Cepstral Coefficients (PNCC), Mel-Frequency Cepstral Coefficients (MFCC) after applying an extended spectral subtraction method, and Sphinx embedded denoising features from recent sphinx versions. Although removing C0 in MFCC-based features leads to a slight decrease in speech recognition performance, it makes the speech recognition system independent on the speech signal level. With multi-condition training, the three sets of noise-robust features lead to a rather similar behavior of performance with respect to SNR and presence of non-speech data. Overall, best performance is achieved with the extended spectral subtraction approach. Also, the performance of the PNCC features appears to be dependent on the initialization of the normalization factor.
Keywords
"Speech","Noise measurement","Mel frequency cepstral coefficient","Speech recognition","Training","Signal to noise ratio","Hidden Markov models"
Publisher
ieee
Conference_Titel
Signal Processing Conference (EUSIPCO), 2015 23rd European
Electronic_ISBN
2076-1465
Type
conf
DOI
10.1109/EUSIPCO.2015.7362617
Filename
7362617
Link To Document