• DocumentCode
    3495136
  • Title

    The SuperSID project: exploiting high-level information for high-accuracy speaker recognition

  • Author

    Reynolds, Douglas ; Andrews, Walter ; Campbell, Joseph ; Navratil, Jiri ; Peskin, Barbara ; Adami, Andrea ; Jin, Qin ; Klusacek, Dalibor ; Abramson, Joy ; Mihaescu, Radu ; Godfrey, Jack ; Jones, Doug ; Xiang, Sing

  • Volume
    4
  • fYear
    2003
  • fDate
    6-10 April 2003
  • Abstract
    The area of automatic speaker recognition has been dominated by systems using only short-term, low-level acoustic information, such as cepstral features. While these systems have indeed produced very low error rates, they ignore other levels of information beyond low-level acoustics that convey speaker information. Recently published work has shown examples that such high-level information can be used successfully in automatic speaker recognition systems and has the potential to improve accuracy and add robustness. For the 2002 JHU CLSP summer workshop, the SuperSID project (http://www.clsp.jhu.edu/ws2002/groups/supersid/) was undertaken to exploit these high-level information sources and dramatically increase speaker recognition accuracy on a defined NIST evaluation corpus and task. The paper provides an overview of the structure, data, task, tools, and accomplishments of this project. Wide ranging approaches using pronunciation models, prosodic dynamics, pitch and duration features, phone streams, and conversational interactions were explored and developed. We show how these novel features and classifiers indeed provide complementary information and can be fused together to drive down the equal error rate on the 2001 NIST extended data task to 0.2% - a 71% relative reduction in error over the previous state of the art.
  • Keywords
    acoustic signal processing; cepstral analysis; error statistics; linguistics; natural languages; reviews; speaker recognition; speech processing; GMM cepstra system; NIST evaluation corpus; SuperSID project; automatic speaker recognition accuracy; cepstral features; conversational interactions; duration features; equal error rate; evaluation task; high-level information; low-level acoustic information; phone streams; pitch features; pronunciation models; prosodic dynamics; Cepstral analysis; Data mining; Error analysis; Humans; Lifting equipment; Loudspeakers; NIST; Natural languages; Speaker recognition; Speech recognition;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Acoustics, Speech, and Signal Processing, 2003. Proceedings. (ICASSP '03). 2003 IEEE International Conference on
  • ISSN
    1520-6149
  • Print_ISBN
    0-7803-7663-3
  • Type

    conf

  • DOI
    10.1109/ICASSP.2003.1202760
  • Filename
    1202760