Dynamic compensation of HMM variances using the feature enhancement uncertainty computed from a parametric model of speech distortion

Author

Deng, Li ; Droppo, Jasha ; Acero, Alex

Author_Institution

Microsoft Res., Redmond, WA, USA

Volume

13

Issue

3

fYear

2005

fDate

5/1/2005 12:00:00 AM

Firstpage

412

Lastpage

421

Abstract

This paper presents a new technique for dynamic, frame-by-frame compensation of the Gaussian variances in the hidden Markov model (HMM), exploiting the feature variance or uncertainty estimated during the speech feature enhancement process, to improve noise-robust speech recognition. The new technique provides an alternative to the Bayesian predictive classification decision rule by carrying out an integration over the feature space instead of over the model-parameter space, offering a much simpler system implementation, lower computational cost, and dynamic compensation capabilities at the frame level. The computation of the feature enhancement variances is carried out using a probabilistic and parametric model of speech distortion, free from the use of any stereo training data. Dynamic compensation of the Gaussian variances in the HMM recognizer is derived, which is simply enlarging the HMM Gaussian variances by the feature enhancement variances. Experimental evaluation using the full Aurora2 test data sets demonstrates a significant digit error rate reduction, averaged over all noisy and signal-to-noise-ratio conditions, compared with the baseline that did not exploit the enhancement variance information. When the true enhancement variances are used, further dramatic error rate reduction is observed, indicating the strong potential for the new technique and the strong need for high accuracy in estimating the variances associated with feature enhancement. All the results, using either the true variances of the enhanced features or the estimated ones, show that the greatest contribution to recognizer´s performance improvement is due to the use of the uncertainty for the static features, next due to the delta features, and the least due to the delta-delta features.

Keywords

Bayes methods; distortion; error statistics; hidden Markov models; noise; speech enhancement; speech recognition; Bayesian predictive classification decision rule; Gaussian variance; computational cost; delta feature; delta-delta feature; digit error rate reduction; enhancement variance information; feature space; frame-by-frame compensation; hidden Markov model; noise-robust speech recognition; signal-to-noise-ratio condition; speech distortion; speech feature enhancement process; static feature; stereo training data; system implementation; Bayesian methods; Error analysis; Hidden Markov models; Noise robustness; Parametric statistics; Predictive models; Speech enhancement; Speech processing; Speech recognition; Uncertainty; Dynamic variance compensation; hidden Markov model (HMM) variance; noise-robust automatic speech recognition (ASR); parametric environment model; speech feature enhancement; uncertainty in feature enhancement;

fLanguage

English

Journal_Title

Speech and Audio Processing, IEEE Transactions on

Publisher

ieee

ISSN

1063-6676

Type

jour

DOI

10.1109/TSA.2005.845814

Filename

1420375