DocumentCode :
3087588
Title :
Crim´s French speech transcription system for ETAPE 2011
Author :
Gupta, V. ; Boulianne, Gilles ; Osterrath, Frederic ; Ouellet, Pierre
fYear :
2013
fDate :
12-15 May 2013
Firstpage :
351
Lastpage :
356
Abstract :
This paper describes the French broadcast speech transcription system by CRIM for the ETAPE 2011 evaluation. The key elements in this recognizer include over 140,000-word dictionary, 478 hours of audio for training the acoustic models, feature-space MMI and boosted MMI discriminative training of the acoustic models, variable-frame-rate decoding with trigram language model, lattice rescoring with quadgram language model, soft penalty on silence models, confusion network decoding with minimum Bayes risk, and combining multiple recognizers with ROVER. Recognition enhancements after the ETAPE evaluation include discriminative training of the subspace Gaussian mixture models and lattice rescoring with neural net language models.
Keywords :
Bayes methods; Gaussian processes; decoding; neural nets; speech coding; speech recognition; variable rate codes; ETAPE 2011 evaluation; French broadcast speech transcription system; ROVER; acoustic models; boosted MMI discriminative training; confusion network decoding; feature-space; lattice rescoring; minimum Bayes risk; multiple recognizers; neural net language models; quadgram language model; recognition enhancements; silence models; soft penalty; subspace Gaussian mixture models; trigram language model; variable-frame-rate decoding; Acoustics; Data models; Decoding; Error analysis; Hidden Markov models; Speech; Training;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Systems, Signal Processing and their Applications (WoSSPA), 2013 8th International Workshop on
Conference_Location :
Algiers
Type :
conf
DOI :
10.1109/WoSSPA.2013.6602390
Filename :
6602390
Link To Document :
بازگشت