DocumentCode :
1733996
Title :
WebSum: Enhanced SumBasic algorithm for Web site summarization
Author :
Tee, Jason Yong-Jin ; Soon, Lay-Ki ; Ting, Choo-Yee
Author_Institution :
Fac. of Comput. & Inf., Multimedia Univ. Cyberjaya, Cyberjaya, Malaysia
fYear :
2012
Firstpage :
137
Lastpage :
142
Abstract :
Due to the rapid increase of information in the World Wide Web, there exists an explosion of information on the Web that may overwhelm the common Web user. The Web user may find it quicker or more efficient to browse the Web by reading summaries of Web sites. This paper proposes WebSum to compress Web site content into a summary. WebSum is an enhancement of the SumBasic algorithm, that was mainly used for multi-document summarization. In the case of Web sites, we find that several Web characteristics such as title and keywords can be used to extract sentences that may represent the overall topic of the Web site. Initial results show that WebSum is able to reveal sentences relate to the concept of the Web site. WebSum is then evaluated against the original algorithm of SumBasic.
Keywords :
Web sites; document handling; SumBasic algorithm; Web site content compression; Web site summarization; WebSum; World Wide Web; keywords; multidocument summarization; sentence extraction; title; Computational linguistics; Data mining; Educational institutions; Indexes; Multimedia communication; Web pages;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Data Mining and Optimization (DMO), 2012 4th Conference on
Conference_Location :
Langkawi
Print_ISBN :
978-1-4673-2717-6
Type :
conf
DOI :
10.1109/DMO.2012.6329812
Filename :
6329812
Link To Document :
بازگشت