DocumentCode :
2091490
Title :
Evaluation of TnT Tagger for Spanish
Author :
Carrasco, Raúl Morales ; Gelbukh, Alexander
Author_Institution :
Instituto Tecnologico de Puebla, Mexico
fYear :
2003
fDate :
8-12 Sept. 2003
Firstpage :
18
Lastpage :
25
Abstract :
Part of speech (POS) tagger is a necessary module in many natural language text processing tasks. A POS tagger is a program that accepts an unprepared raw text in input and to each word adds a tag specifying its grammatical properties, such as part of speech, number, person, etc. One of popular POS taggers - TnT tagger - has been extensively tested for English and some other languages. This paper reports on its evaluation for Spanish language. Error analysis is reported, explaining how some specific features of Spanish language affect tagger performance. It is reported that on Spanish texts TnT shows overall tagging accuracy between 92.5% and 95.84%, specifically, between 95.47% and 98.56% on known words and between 75.57% and 83.49% on unknown words. Results show that TnT has reached a good level of maturity and is helpful enough for NLP tasks.
Keywords :
computational linguistics; grammars; natural languages; text analysis; English; NLP tasks; Spanish; TnT tagger; error analysis; natural language processing; part of speech tagger; text processing; Character recognition; Error analysis; Mood; Natural languages; Speech processing; Speech recognition; Tagging; Testing; Text processing; Text recognition;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Computer Science, 2003. ENC 2003. Proceedings of the Fourth Mexican International Conference on
Print_ISBN :
0-7695-1915-6
Type :
conf
DOI :
10.1109/ENC.2003.1232869
Filename :
1232869
Link To Document :
بازگشت