• DocumentCode
    3527300
  • Title

    Single-channel speech separation and recognition using loopy belief propagation

  • Author

    Rennie, Steven J. ; Hershey, John R. ; Olsen, Peder A.

  • Author_Institution
    IBM T.J. Watson Res. Center, Yorktown Heights, NY
  • fYear
    2009
  • fDate
    19-24 April 2009
  • Firstpage
    3845
  • Lastpage
    3848
  • Abstract
    We address the problem of single-channel speech separation and recognition using loopy belief propagation in a way that enables efficient inference for an arbitrary number of speech sources. The graphical model consists of a set of N Markov chains, each of which represents a language model or grammar for a given speaker. A Gaussian mixture model with shared states is used to model the hidden acoustic signal for each grammar state of each source. The combination of sources is modeled in the log spectrum domain using non-linear interaction functions. Previously, temporal inference in such a model has been performed using an N-dimensional Viterbi algorithm that scales exponentially with the number of sources. In this paper, we describe a loopy message passing algorithm that scales linearly with language model size. The algorithm achieves human levels of performance, and is an order of magnitude faster than competitive systems for two speakers.
  • Keywords
    Gaussian processes; Markov processes; belief maintenance; speech recognition; Gaussian mixture model; Markov chains; Viterbi algorithm; hidden acoustic signal; log spectrum domain; loopy belief propagation; loopy message passing; non-linear interaction functions; single-channel speech separation; speech recognition; temporal inference; Automatic speech recognition; Belief propagation; Computational efficiency; Graphical models; Hidden Markov models; Humans; Inference algorithms; Loudspeakers; Speech recognition; Viterbi algorithm; ASR; Algonquin; Iroquois; Max model; Speech separation; factorial hidden Markov models; loopy belief propagation;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Acoustics, Speech and Signal Processing, 2009. ICASSP 2009. IEEE International Conference on
  • Conference_Location
    Taipei
  • ISSN
    1520-6149
  • Print_ISBN
    978-1-4244-2353-8
  • Electronic_ISBN
    1520-6149
  • Type

    conf

  • DOI
    10.1109/ICASSP.2009.4960466
  • Filename
    4960466