DocumentCode :
3428234
Title :
Spoken term detection for Turkish Broadcast News
Author :
Parlak, Siddika ; Saraçlar, Murat
Author_Institution :
Dept. of Electr. & Electron. Eng., Bogazici Univ., Istanbul
fYear :
2008
fDate :
March 31 2008-April 4 2008
Firstpage :
5244
Lastpage :
5247
Abstract :
In this paper, we present a baseline spoken term detection (STD) system for Turkish broadcast news. The agglutinative structure of Turkish causes a high out-of-vocabulary (OOV) rate and increases word error rate (WER) in automatic speech recognition. Several approaches are attempted to reduce this negative effect on the STD system. Sub-word units are used to handle the OOV queries and lattice-based indexing is used to obtain different operating points and handle high WER cases. A recently proposed method for setting term specific thresholds is also evaluated and extended to allow us to choose an operating point suitable for our needs. Best results are obtained by using a cascade of word and sub-word lattice indices with term-thresholding.
Keywords :
indexing; information retrieval; speech recognition; text analysis; word processing; Turkish broadcast news; automatic speech recognition; lattice-based indexing; out-of-vocabulary queries; spoken term detection; sub-word lattice indices; term-thresholding; word error rate; Automatic speech recognition; Broadcasting; Design methodology; Error analysis; Handicapped aids; Indexing; Information retrieval; Lattices; Speech processing; Speech recognition; audio indexing; information retrieval; speech recognition; spoken term detection;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Acoustics, Speech and Signal Processing, 2008. ICASSP 2008. IEEE International Conference on
Conference_Location :
Las Vegas, NV
ISSN :
1520-6149
Print_ISBN :
978-1-4244-1483-3
Electronic_ISBN :
1520-6149
Type :
conf
DOI :
10.1109/ICASSP.2008.4518842
Filename :
4518842
Link To Document :
بازگشت