Crim´s French speech transcription system for ETAPE 2011

Author

Gupta, V. ; Boulianne, Gilles ; Osterrath, Frederic ; Ouellet, Pierre

fYear

2013

fDate

12-15 May 2013

Firstpage

351

Lastpage

356

Abstract

This paper describes the French broadcast speech transcription system by CRIM for the ETAPE 2011 evaluation. The key elements in this recognizer include over 140,000-word dictionary, 478 hours of audio for training the acoustic models, feature-space MMI and boosted MMI discriminative training of the acoustic models, variable-frame-rate decoding with trigram language model, lattice rescoring with quadgram language model, soft penalty on silence models, confusion network decoding with minimum Bayes risk, and combining multiple recognizers with ROVER. Recognition enhancements after the ETAPE evaluation include discriminative training of the subspace Gaussian mixture models and lattice rescoring with neural net language models.

Keywords

Bayes methods; Gaussian processes; decoding; neural nets; speech coding; speech recognition; variable rate codes; ETAPE 2011 evaluation; French broadcast speech transcription system; ROVER; acoustic models; boosted MMI discriminative training; confusion network decoding; feature-space; lattice rescoring; minimum Bayes risk; multiple recognizers; neural net language models; quadgram language model; recognition enhancements; silence models; soft penalty; subspace Gaussian mixture models; trigram language model; variable-frame-rate decoding; Acoustics; Data models; Decoding; Error analysis; Hidden Markov models; Speech; Training;

fLanguage

English

Publisher

ieee

Conference_Titel

Systems, Signal Processing and their Applications (WoSSPA), 2013 8th International Workshop on

Conference_Location

Algiers

Type

conf

DOI

10.1109/WoSSPA.2013.6602390

Filename

6602390