• DocumentCode
    1060222
  • Title

    Soft Mask Methods for Single-Channel Speaker Separation

  • Author

    Reddy, Aarthi M. ; Raj, Bhiksha

  • Author_Institution
    McGill Univ., Montreal
  • Volume
    15
  • Issue
    6
  • fYear
    2007
  • Firstpage
    1766
  • Lastpage
    1776
  • Abstract
    The problem of single-channel speaker separation attempts to extract a speech signal uttered by the speaker of interest from a signal containing a mixture of acoustic signals. Most algorithms that deal with this problem are based on masking, wherein unreliable frequency components from the mixed signal spectrogram are suppressed, and the reliable components are inverted to obtain the speech signal from speaker of interest. Most current techniques estimate this mask in a binary fashion, resulting in a hard mask. In this paper, we present two techniques to separate out the speech signal of the speaker of interest from a mixture of speech signals. One technique estimates all the spectral components of the desired speaker. The second technique estimates a soft mask that weights the frequency subbands of the mixed signal. In both cases, the speech signal of the speaker of interest is reconstructed from the complete spectral descriptions obtained. In their native form, these algorithms are computationally expensive. We also present fast factored approximations to the algorithms. Experiments reveal that the proposed algorithms can result in significant enhancement of individual speakers in mixed recordings, consistently achieving better performance than that obtained with hard binary masks.
  • Keywords
    speaker recognition; speech enhancement; speech intelligibility; signal spectrogram; single-channel speaker separation; soft mask methods; speakers enhancement; speech signal extraction; Computational modeling; Ear; Frequency estimation; Humans; Loudspeakers; Signal processing; Source separation; Spectrogram; Speech coding; Speech recognition; Signal separation; soft masks; speaker separation;
  • fLanguage
    English
  • Journal_Title
    Audio, Speech, and Language Processing, IEEE Transactions on
  • Publisher
    ieee
  • ISSN
    1558-7916
  • Type

    jour

  • DOI
    10.1109/TASL.2007.901310
  • Filename
    4276763