Title :
Enhanced Sparse Imputation Techniques for a Robust Speech Recognition Front-End
Author :
Qun Feng Tan ; Georgiou, P.G. ; Narayanan, S.
Author_Institution :
Dept. of Electr. Eng., Univ. of Southern California, Los Angeles, CA, USA
Abstract :
Missing data techniques (MDTs) have been widely employed and shown to improve speech recognition results under noisy conditions. This paper presents a new technique which improves upon previously proposed sparse imputation techniques relying on the least absolute shrinkage and selection operator (LASSO). LASSO is widely employed in compressive sensing problems. However, the problem with LASSO is that it does not satisfy oracle properties in the event of a highly collinear dictionary, which happens with features extracted from most speech corpora. When we say that a variable selection procedure satisfies the oracle properties, we mean that it enjoys the same performance as though the underlying true model is known. Through experiments on the Aurora 2.0 noisy spoken digits database, we demonstrate that the Least Angle Regression implementation of the Elastic Net (LARS-EN) algorithm is able to better exploit the properties of a collinear dictionary, and thus is significantly more robust in terms of basis selection when compared to LASSO on the continuous digit recognition task with estimated mask. In addition, we investigate the effects and benefits of a good measure of sparsity on speech recognition rates. In particular, we demonstrate that a good measure of sparsity greatly improves speech recognition rates, and that the LARS modification of LASSO and LARS-EN can be terminated early to achieve improved recognition results, even though the estimation error is increased.
Keywords :
feature extraction; speech recognition; Aurora 2.0 noisy spoken digits database; LASSO; collinear dictionary; compressive sensing problems; continuous digit recognition task; enhanced sparse imputation techniques; estimation error; feature extraction; least absolute shrinkage and selection operator; least angle regression implementation of the elastic net; missing data techniques; robust speech recognition front-end; Dictionaries; Noise measurement; Optimization; Signal to noise ratio; Speech; Speech recognition; Automatic speech recognition (ASR); compressive sensing; convex optimization; missing data techniques (MDTs); robustness; sparse representation;
Journal_Title :
Audio, Speech, and Language Processing, IEEE Transactions on
DOI :
10.1109/TASL.2011.2136337