مرکز منطقه ای اطلاع رساني علوم و فناوري - Capturing Complementary Information via Reversed Filter Bank and Parallel Implementation with MFCC for Improved Text-Independent Speaker Identification

DocumentCode :

1925730

Title :

Capturing Complementary Information via Reversed Filter Bank and Parallel Implementation with MFCC for Improved Text-Independent Speaker Identification

Author :

Chakroborty, Sandipan ; Roy, Anindya ; Majumdar, Sourav ; Saha, Goutam

Author_Institution :

Dept. of Electron. and Electr. Commun. Eng., Indian Inst. of Technol., Kharagpur

fYear :

2007

fDate :

5-7 March 2007

Firstpage :

463

Lastpage :

467

Abstract :

A state of the art speaker identification (SI) system requires a robust feature extraction unit followed by a speaker modeling scheme for generalized representation of these features. Over the years, mel-frequency cepstral coefficients (MFCC) modeled on the human auditory system have been used as a standard acoustic feature set for SI applications. However, due to the structure of its filter bank, it captures vocal tract characteristics more effectively in the lower frequency regions. This work proposes a new set of features using a complementary filter bank structure which improves distinguishability of speaker specific cues present in the higher frequency zone. Unlike high level features that are difficult to extract, the proposed feature set involves little computational burden during the extraction process. When combined with MFCC via a parallel implementation of speaker models, the proposed feature improves performance baseline of MFCC based system. The proposition is validated by experiments conducted on two different kinds of databases namely YOHO (microphone speech) and POLYCOST (telephone speech) with two different classifier paradigms, namely Gaussian Mixture Models (GMM) and Polynomial Classifier (PC) and for various model orders

Keywords :

audio databases; audio signal processing; channel bank filters; feature extraction; pattern classification; speaker recognition; Gaussian mixture models; MFCC; POLYCOST database; YOHO database; classifier paradigms; complementary information; human auditory system; mel-frequency cepstral coefficients; parallel implementation; polynomial classifier; reversed filter bank; robust feature extraction unit; speaker modeling scheme; standard acoustic feature set; text-independent speaker identification system; vocal tract characteristics; Acoustic applications; Auditory system; Cepstral analysis; Feature extraction; Filter bank; Humans; Loudspeakers; Mel frequency cepstral coefficient; Robustness; Speech;

fLanguage :

English

Publisher :

ieee

Conference_Titel :

Computing: Theory and Applications, 2007. ICCTA '07. International Conference on

Conference_Location :

Kolkata

Print_ISBN :

0-7695-2770-1

Type :

conf

DOI :

10.1109/ICCTA.2007.35

Filename :

4127413

Link To Document :

https://search.ricest.ac.ir/dl/search/defaultta.aspx?DTC=49&DC=1925730