• DocumentCode
    1518811
  • Title

    Exemplar-Based Sparse Representation Features: From TIMIT to LVCSR

  • Author

    Sainath, Tara N. ; Ramabhadran, Bhuvana ; Picheny, Michael ; Nahamoo, David ; Kanevsky, Dimitri

  • Author_Institution
    IBM T. J. Watson Res. Center, Yorktown Heights, NY, USA
  • Volume
    19
  • Issue
    8
  • fYear
    2011
  • Firstpage
    2598
  • Lastpage
    2613
  • Abstract
    The use of exemplar-based methods, such as support vector machines (SVMs), k-nearest neighbors (kNNs) and sparse representations (SRs), in speech recognition has thus far been limited. Exemplar-based techniques utilize information about individual training examples and are computationally expensive, making it particularly difficult to investigate these methods on large-vocabulary continuous speech recognition (LVCSR) tasks. While research in LVCSR provides a good testbed to tackle real-world speech recognition problems, research in this area suffers from two main drawbacks. First, the overall complexity of an LVCSR system makes error analysis quite difficult. Second, exploring new research ideas on LVCSR tasks involves training and testing state-of-the-art LVCSR systems, which can render a large turnaround time. This makes a small vocabulary task such as TIMIT more appealing. TIMIT provides a phonetically rich and hand-labeled corpus that allows easy insight into new algorithms. However, research ideas explored for small vocabulary tasks do not always provide gains on LVCSR systems. In this paper, we combine the advantages of using both small and large vocabulary tasks by taking well-established techniques used in LVCSR systems and applying them on TIMIT to establish a new baseline. We then utilize these existing LVCSR techniques in creating a novel set of exemplar-based sparse representation (SR) features. Using these existing LVCSR techniques, we achieve a phonetic error rate (PER) of 19.4% on the TIMIT task. The additional use of SR features reduce the PER to 18.6%. We then explore applying the SR features to a large vocabulary Broadcast News task, where we achieve a 0.3% absolute reduction in word error rate (WER).
  • Keywords
    error analysis; error statistics; speech recognition; support vector machines; SR feature; TIMIT; error analysis; exemplar-based sparse representation feature; exemplar-based technique; hand-labeled corpus; k-nearest neighbor; large vocabulary broadcast news task; large vocabulary continuous speech recognition; phonetic error rate; real world speech recognition problem; state-of-the-art LVCSR system complexity; support vector machine; vocabulary task; word error rate reduction; Error analysis; Hidden Markov models; Speech recognition; Training; Vocabulary; Exemplar-based techniques; sparse representations (SRs); speech recognition;
  • fLanguage
    English
  • Journal_Title
    Audio, Speech, and Language Processing, IEEE Transactions on
  • Publisher
    ieee
  • ISSN
    1558-7916
  • Type

    jour

  • DOI
    10.1109/TASL.2011.2155060
  • Filename
    5770195