Title :
Crim´s French speech transcription system for ETAPE 2011
Author :
Gupta, V. ; Boulianne, Gilles ; Osterrath, Frederic ; Ouellet, Pierre
Abstract :
This paper describes the French broadcast speech transcription system by CRIM for the ETAPE 2011 evaluation. The key elements in this recognizer include over 140,000-word dictionary, 478 hours of audio for training the acoustic models, feature-space MMI and boosted MMI discriminative training of the acoustic models, variable-frame-rate decoding with trigram language model, lattice rescoring with quadgram language model, soft penalty on silence models, confusion network decoding with minimum Bayes risk, and combining multiple recognizers with ROVER. Recognition enhancements after the ETAPE evaluation include discriminative training of the subspace Gaussian mixture models and lattice rescoring with neural net language models.
Keywords :
Bayes methods; Gaussian processes; decoding; neural nets; speech coding; speech recognition; variable rate codes; ETAPE 2011 evaluation; French broadcast speech transcription system; ROVER; acoustic models; boosted MMI discriminative training; confusion network decoding; feature-space; lattice rescoring; minimum Bayes risk; multiple recognizers; neural net language models; quadgram language model; recognition enhancements; silence models; soft penalty; subspace Gaussian mixture models; trigram language model; variable-frame-rate decoding; Acoustics; Data models; Decoding; Error analysis; Hidden Markov models; Speech; Training;
Conference_Titel :
Systems, Signal Processing and their Applications (WoSSPA), 2013 8th International Workshop on
Conference_Location :
Algiers
DOI :
10.1109/WoSSPA.2013.6602390