• DocumentCode
    1245780
  • Title

    Robust text-independent speaker identification using Gaussian mixture speaker models

  • Author

    Reynolds, Douglas A. ; Rose, Richard C.

  • Author_Institution
    Lincoln Lab., MIT, Lexington, MA, USA
  • Volume
    3
  • Issue
    1
  • fYear
    1995
  • fDate
    1/1/1995 12:00:00 AM
  • Firstpage
    72
  • Lastpage
    83
  • Abstract
    This paper introduces and motivates the use of Gaussian mixture models (GMM) for robust text-independent speaker identification. The individual Gaussian components of a GMM are shown to represent some general speaker-dependent spectral shapes that are effective for modeling speaker identity. The focus of this work is on applications which require high identification rates using short utterance from unconstrained conversational speech and robustness to degradations produced by transmission over a telephone channel. A complete experimental evaluation of the Gaussian mixture speaker model is conducted on a 49 speaker, conversational telephone speech database. The experiments examine algorithmic issues (initialization, variance limiting, model order selection), spectral variability robustness techniques, large population performance, and comparisons to other speaker modeling techniques (uni-modal Gaussian, VQ codebook, tied Gaussian mixture, and radial basis functions). The Gaussian mixture speaker model attains 96.8% identification accuracy using 5 second clean speech utterances and 80.8% accuracy using 15 second telephone speech utterances with a 49 speaker population and is shown to outperform the other speaker modeling techniques on an identical 16 speaker telephone speech task
  • Keywords
    Gaussian processes; speaker recognition; speech processing; telephony; Gaussian mixture speaker models; VQ codebook; algorithmic issues; conversational telephone speech database; degradations; high identification rates; identification accuracy; initialization; large population performance; model order selection; modeling; radial basis functions; robust text-independent speaker identification; short utterance; speaker-dependent spectral shapes; spectral variability robustness techniques; telephone channel; tied Gaussian mixture; transmission; unconstrained conversational speech; uni-modal Gaussian technique; variance limiting; Data mining; Databases; Degradation; Helium; Loudspeakers; Robustness; Speaker recognition; Spectral shape; Speech analysis; Telephony;
  • fLanguage
    English
  • Journal_Title
    Speech and Audio Processing, IEEE Transactions on
  • Publisher
    ieee
  • ISSN
    1063-6676
  • Type

    jour

  • DOI
    10.1109/89.365379
  • Filename
    365379