Title of article
Feature reduction techniques for Arabic text categorization
Author/Authors
Rehab Duwairi1، نويسنده , , Mohammad Nayef Al-Refai2، نويسنده , , Natheer Khasawneh، نويسنده ,
Issue Information
ماهنامه با شماره پیاپی سال 2009
Pages
6
From page
2347
To page
2352
Abstract
This paper presents and compares three feature reduction techniques that were applied to Arabic text. The techniques include stemming, light stemming, and word clusters. The effects of the aforementioned techniques were studied and analyzed on the K-nearest-neighbor classifier. Stemming reduces words to their stems. Light stemming, by comparison, removes common affixes from words without reducing them to their stems. Word clusters group synonymous words into clusters and each cluster is represented by a single word. The purpose of employing the previous methods is to reduce the size of document vectors without affecting the accuracy of the classifiers. The comparison metric includes size of document vectors, classification time, and accuracy (in terms of precision and recall). Several experiments were carried out using four different representations of the same corpus: the first version uses stem-vectors, the second uses light stem-vectors, the third uses word clusters, and the fourth uses the original words (without any transformation) as representatives of documents. The corpus consists of 15,000 documents that fall into three categories: sports, economics, and politics. In terms of vector sizes and classification time, the stemmed vectors consumed the smallest size and the least time necessary to classify a testing dataset that consists of 6,000 documents. The light stemmed vectors superseded the other three representations in terms of classification accuracy.
Journal title
Journal of the American Society for Information Science and Technology
Serial Year
2009
Journal title
Journal of the American Society for Information Science and Technology
Record number
994093
Link To Document