Title of article :
Robust recognition of childrens speech
Author/Authors :
A.، Potamianos, نويسنده , , S.، Narayanan, نويسنده ,
Issue Information :
روزنامه با شماره پیاپی سال 2003
Pages :
-602
From page :
603
To page :
0
Abstract :
Developmental changes in speech production introduce age-dependent spectral and temporal variability in the speech signal produced by children. Such variabilities pose challenges for robust automatic recognition of childrenʹs speech. Through an analysis of age-related acoustic characteristics of childrenʹs speech in the context of automatic speech recognition (ASR), effects such as frequency scaling of spectral envelope parameters are demonstrated. Recognition experiments using acoustic models trained from adult speech and tested against speech from children of various ages clearly show performance degradation with decreasing age. On average, the word error rates are two to five times worse for children speech than for adult speech. Various techniques for improving ASR performance on childrenʹs speech are reported. A speaker normalization algorithm that combines frequency warping and model transformation is shown to reduce acoustic variability and significantly improve ASR performance for children speakers (by 25-45% under various model training and testing conditions). The use of age-dependent acoustic models further reduces word error rate by 10%. The potential of using piece-wise linear and phoneme-dependent frequency warping algorithms for reducing the variability in the acoustic feature space of children is also investigated.
Keywords :
waveguide transition , Laminated waveguide , low-temperature co-fired ceramic (LTCC) , rectangular waveguide (RWG) , millimeter wave
Journal title :
IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING
Serial Year :
2003
Journal title :
IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING
Record number :
86935
Link To Document :
بازگشت