• DocumentCode
    172543
  • Title

    Effectiveness of multiscale fractal dimension-based phonetic segmentation in speech synthesis for low resource language

  • Author

    Zaki, Mohammadi ; Shah, J. Nirmesh ; Patil, Hemant A.

  • Author_Institution
    Dhirubhai Ambani Inst. of Inf. & Commun. Technol. (DA-IICT), Gandhinagar, India
  • fYear
    2014
  • fDate
    20-22 Oct. 2014
  • Firstpage
    103
  • Lastpage
    106
  • Abstract
    Phonetic segmentation plays a key role in developing various speech applications. In this work, we propose to use various features for automatic phonetic segmentation task for forced Viterbi alignment and compare their effectiveness. We propose to use novel multiscale fractal dimension-based features concatenated with Mel-Frequency Cepstral Coefficients (MFCC). The novel features are expected to capture additional nonlinearities in speech production which should improve the performance of segmentation task. However, to evaluate effectiveness of these segmentation algorithms, we require manual accurate phoneme-level labeled data which is not available for low resource languages such as Gujarati (a low resource language and one of the official languages of India). In order to measure effectiveness of various segmentation algorithms, HMM-based speech synthesis system (HTS) for Gujarati have been built. From the subjective and objective evaluations, it is observed that FD-based features for segmentation work moderately better than other state-of-the-art features such as MFCC, Perceptual Linear Prediction Cepstral Coefficients (PLP-CC), Cochlear Filter Cepstral Coefficients (CFCC), and RelAtive SpecTrAl (RASTA)-based PLP-CC. The Mean Opinion Score (MOS) and the Degraded-MOS, which are the measures of naturalness indicate an improvement of 9.69% with the proposed features from the MFCC (which is found to be the best among the other features) based features.
  • Keywords
    cepstral analysis; natural language processing; speech synthesis; Gujarati; HMM; HTS; MFCC; automatic phonetic segmentation task; degraded-MOS; forced Viterbi alignment; low resource language; mean opinion score; mel-frequency cepstral coefficients; multiscale fractal dimension-based phonetic segmentation; objective evaluation; performance improvement; phoneme-level labeled data; segmentation work; speech production; speech synthesis system; subjective evaluation; Fractals; Hidden Markov models; High-temperature superconductors; Mel frequency cepstral coefficient; Production; Speech; Speech synthesis; HMM, HTS; multifractal; multiscale fractal dimension;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Asian Language Processing (IALP), 2014 International Conference on
  • Conference_Location
    Kuching
  • Type

    conf

  • DOI
    10.1109/IALP.2014.6973508
  • Filename
    6973508