DocumentCode :
2273079
Title :
Table Based Single Pass Algorithm for Clustering Electronic Documents in 20NewsGroups
Author :
Jo, Taeho ; Jo, Geun-Sik
Author_Institution :
Sch. of Comput. & Inf. Eng., Inha Univ., Incheon
fYear :
2008
fDate :
10-11 July 2008
Firstpage :
66
Lastpage :
71
Abstract :
This research proposes a modified version of single pass algorithm which is specialized for text clustering. Encoding documents into numerical vectors for using the traditional version of single pass algorithm causes the two main problems: huge dimensionality and sparse distribution. Therefore, in order to address the two problems, this research modifies the single pass algorithm into its version where documents are encoded into not numerical vectors but alternative forms. In the proposed version, documents are mapped into tables and a similarity of two documents is computed by comparing their tables with each other. The goal of this research is to improve the performance of single pass algorithm for text clustering by modifying it into the specialized version.
Keywords :
document handling; information resources; pattern clustering; text analysis; document encoding; electronic document clustering; newsgroups; single pass algorithm; table based single pass algorithm; text clustering; Application software; Clustering algorithms; Computer applications; Conferences; Encoding; Kernel; Support vector machines; Text categorization; Text mining; Unsupervised learning;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Semantic Computing and Applications, 2008. IWSCA '08. IEEE International Workshop on
Conference_Location :
Incheon
Print_ISBN :
978-0-7695-3317-9
Electronic_ISBN :
978-0-7695-3317-9
Type :
conf
DOI :
10.1109/IWSCA.2008.32
Filename :
4573152
Link To Document :
بازگشت