مرکز منطقه ای اطلاع رساني علوم و فناوري - Part-of-speech labeling for Reuters database

DocumentCode :

3686547

Title :

Part-of-speech labeling for Reuters database

Author :

R. Cretulescu;A. David;D. Morariu;L. Vintan

Author_Institution :

Comput. Sci. &

fYear :

2015

Firstpage :

117

Lastpage :

122

Abstract :

Even if the Vector Space Model used for document representation in information retrieval systems integrates a small quantity of knowledge it continues to be used due to its computational cost, speed execution and simplicity. We try to improve this document representation by adding some syntactic information such as the parts of speech. In this paper, we have evaluated three different tagging algorithms in order to select the most suitable tagger for using it to tag the Reuters dataset. In this work, we have evaluated the taggers using only five different parts of speech: noun, verb, adverb, adjective and others. We considered these particular tags being the most representative for describing the documents into these parts of speech space.

Keywords :

"Speech","Tagging","Training","Accuracy","Labeling","Testing","Syntactics"

Publisher :

ieee

Conference_Titel :

System Theory, Control and Computing (ICSTCC), 2015 19th International Conference on

Type :

conf

DOI :

10.1109/ICSTCC.2015.7321279

Filename :

7321279

Link To Document :

https://search.ricest.ac.ir/dl/search/defaultta.aspx?DTC=49&DC=3686547