DocumentCode :
1984703
Title :
Evaluation of statistical part of speech tagging of persian text
Author :
Tasharofi, Samira ; Raja, F. ; Oroumchian, F. ; Rahgozar, Masoud
Author_Institution :
Electr. & Comput. Eng. Dept., Univ. of Tehran, Tehran
fYear :
2007
fDate :
12-15 Feb. 2007
Firstpage :
1
Lastpage :
4
Abstract :
Part of Speech (POS) tagging is an essential part of text processing applications. A POS tagger assigns a tag to each word of its input text specifying its grammatical properties. One of the popular POS taggers is TnT tagger which was shown to have high accuracy in English and some other languages. It is always interesting to see how a method in one language performs on another language because it would give us insight into the difference and similarities of the languages. In case of statistical methods such as TnT, this will have an added practical advantages also. This paper presents creation of a POS tagged corpus and evaluation of TnT tagger on Persian text. The results of experiments on Persian text show that TnT provides overall tagging accuracy of 96.64%, specifically, 97.01% on known words and 77.77% on unknown words.
Keywords :
natural language processing; speech processing; statistical analysis; text analysis; word processing; English; Persian text; grammatical properties; part of speech; speech tagging; statistical methods; text processing; Application software; Interpolation; Natural languages; Smoothing methods; Speech analysis; Speech processing; Statistical analysis; Tagging; Testing; Text processing;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Signal Processing and Its Applications, 2007. ISSPA 2007. 9th International Symposium on
Conference_Location :
Sharjah
Print_ISBN :
978-1-4244-0778-1
Electronic_ISBN :
978-1-4244-1779-8
Type :
conf
DOI :
10.1109/ISSPA.2007.4555312
Filename :
4555312
Link To Document :
بازگشت