• DocumentCode
    1687545
  • Title

    Effectiveness of discriminative training and feature transformation for reverberated and noisy speech

  • Author

    Tachioka, Yuuki ; Watanabe, Shigetaka ; Hershey, John R.

  • Author_Institution
    Inf. Technol. R&D Center, Mitsubishi Electr., Kamakura, Japan
  • fYear
    2013
  • Firstpage
    6935
  • Lastpage
    6939
  • Abstract
    Automatic speech recognition in the presence of non-stationary interference and reverberation remains a challenging problem. The 2nd `CHiME´ Speech Separation and Recognition Challenge introduces a new and difficult task with time-varying reverberation and non-stationary interference including natural background speech, home noises, or music. This paper establishes baselines using state-of-the-art ASR techniques such as discriminative training and various feature transformation on the middle-vocabulary sub-task of this challenge. In addition, we propose an augmented discriminative feature transformation that introduces arbitrary features to a discriminative feature transformation. We present experimental results showing that discriminative training of model parameters and feature transforms is highly effective for this task, and that the augmented feature transformation provides some preliminary benefits. The training code will be released as an advanced ASR baseline.
  • Keywords
    learning (artificial intelligence); speech recognition; training; transforms; ASR techniques; CHiME; augmented discriminative feature transformation; automatic speech recognition; discriminative training; feature transforms; home noises; middle-vocabulary sub-task; model parameters; music; natural background speech; noisy speech; nonstationary interference; recognition challenge; reverberated speech; speech separation; time-varying reverberation; training code; Hidden Markov models; Mel frequency cepstral coefficient; Noise; Noise measurement; Speech; Speech recognition; Training; Augmented discriminative feature transformation; CHiME challenge; Discriminative training; Feature transformation; Kaldi;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Acoustics, Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on
  • Conference_Location
    Vancouver, BC
  • ISSN
    1520-6149
  • Type

    conf

  • DOI
    10.1109/ICASSP.2013.6639006
  • Filename
    6639006