Automatic speaker independent alignment of continuous speech with its phonetic transcription using a hidden Markov model

Author

Brümmer, JNL ; Boetzer, M.W.

Author_Institution

Dept. of Electr. & Electron. Eng., Stellenbosch Univ., South Africa

fYear

1988

fDate

32318

Firstpage

35

Lastpage

40

Abstract

A way is presented to time-aligned phonetic transcriptions with an acoustic speech waveform using hidden Markov models defined by the transcriptions. Given an utterance of speech and its phonetic transcription, the algorithm will yield the starting and ending times of all the phonemes in the transcription, relative to the start of the utterance. The probabilities for the model are obtained from phoneme duration probabilities and feature probabilities for a few coarse phoneme classes. Because of the coarse classes, the method is speaker-independent. The alignment is accomplished using the Viterbi algorithm. An efficient way of implementing the Viterbi algorithm is given. By using single word transcriptions, the method can be used to detect words in continuous speech, which allows words to be searched for using only their phonetic representations. Two different hidden Markov models (HMM) were used, one with discrete observation symbols and one with continuous observation vectors. The continuous model works better, but the discrete one works faster

Keywords

Markov processes; speech analysis and processing; Viterbi algorithm; acoustic speech waveform; continuous observation vectors; continuous speech; discrete observation symbols; hidden Markov model; phonemes; phonetic transcription; speaker-independent; transcriptions; utterance; Acoustic signal detection; Acoustic waves; Acoustical engineering; Hidden Markov models; Humans; Loudspeakers; Quantization; Speech processing; Speech recognition; Viterbi algorithm;

fLanguage

English

Publisher

ieee

Conference_Titel

Communications and Signal Processing, 1988. Proceedings., COMSIG 88. Southern African Conference on

Conference_Location

Pretoria

Print_ISBN

0-87942-709-4

Type

conf

DOI

10.1109/COMSIG.1988.49298

Filename

49298