• DocumentCode
    3124256
  • Title

    Efficient feature extraction of speaker identification using phoneme mean F-ratio for Chinese

  • Author

    Chen Zhao ; Hongcui Wang ; Songgun Hyon ; Jianguo Wei ; Jianwu Dang

  • Author_Institution
    Sch. of Comput. Sci. & Technol., Tianjin Univ., Tianjin, China
  • fYear
    2012
  • fDate
    5-8 Dec. 2012
  • Firstpage
    345
  • Lastpage
    348
  • Abstract
    The features used for speaker recognition should have more speaker individual information while attenuating the linguistic information. In order to discard the linguistic information effectively, in this paper, we employed the phoneme mean F-ratio method to investigate the different contributions of different frequency region from the point of view of Chinese phoneme, and apply it for speaker identification. It is found that the speaker individual information depending on the phonemes is distributed in different frequency regions of speech sound. Based on the contribution rate, we extracted the new features and combined with GMM model. The experiment for speaker identification task is conducted with a King-ASR Chinese database. Compared with the MFCC feature, the identification error rate with the proposed feature was reduced by 32.94%. The results confirmed that the efficiency of the phoneme mean F-ratio method for improving speaker recognition performance for Chinese.
  • Keywords
    Gaussian processes; cepstral analysis; computational linguistics; feature extraction; natural language processing; speaker recognition; speech processing; Chinese phoneme; GMM model; MFCC feature; contribution rate; feature extraction; frequency regions; identification error rate; king-ASR Chinese database; linguistic information; phoneme mean F-ratio method; phoneme mean f-ratio; speaker identification; speaker individual information; speaker recognition performance; speech sound; Feature extraction; Frequency domain analysis; Hidden Markov models; Mel frequency cepstral coefficient; Speaker recognition; Speech; feature extraction; phoneme mean F-ratio; speaker identification;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Chinese Spoken Language Processing (ISCSLP), 2012 8th International Symposium on
  • Conference_Location
    Kowloon
  • Print_ISBN
    978-1-4673-2506-6
  • Electronic_ISBN
    978-1-4673-2505-9
  • Type

    conf

  • DOI
    10.1109/ISCSLP.2012.6423485
  • Filename
    6423485