DocumentCode :
2228822
Title :
An Efficient Approach for Finding Near Duplicate Web Pages Using Minimum Weight Overlapping Method
Author :
Das, Sankar Narayan ; Mathew, Michael ; Vijayaraghavan, Pramod K.
Author_Institution :
Dept. of Comput. Applic., CUSAT, India
fYear :
2012
fDate :
16-18 April 2012
Firstpage :
121
Lastpage :
126
Abstract :
The existence of billions of web data has severely affected the performance and reliability of web search. The presence of near duplicate web pages plays an important role in this performance degradation while integrating data from heterogeneous sources. Web mining faces huge problems due to the existence of such documents. These pages increase the index storage space and thereby increase the serving cost. By introducing efficient methods to detect and remove such documents from the Web not only decreases the computation time but also increases the relevancy of search results. We aim a novel idea for finding near duplicate web pages which can be incorporated in the field of plagiarism detection, spam detection and focused web crawling scenarios. Here we propose an efficient method for finding near duplicates of an input web page, from a huge repository. A TDW matrix based algorithm is proposed with three phases, rendering, filtering and verification, which receives an input web page and a threshold in its first phase, prefix filtering and positional filtering to reduce the size of record set in the second phase and returns an optimal set of near duplicate web pages in the verification phase by using Minimum Weight Overlapping (MWO) method. The experimental results show that our algorithm outperforms in terms of two benchmark measures, precision and recall, and a reduction in the size of competing record set.
Keywords :
Internet; data mining; document handling; information filtering; rendering (computer graphics); security of data; unsolicited e-mail; TDW matrix based algorithm; Web mining; Web page finding; Web search; document detection; document removal; filtering phase; focused Web crawling scenario; minimum weight overlapping method; near duplicate Web page; plagiarism detection; positional filtering; prefix filtering; rendering phase; spam detection; verification phase; Algorithm design and analysis; Filtering; Rendering (computer graphics); Search engines; Standards; Web pages; Minimum Weight overlapping; Near Duplicate Detection; Positional filtering; Prefix filtering; Term Document Weight Matrix; web page classification;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Information Technology: New Generations (ITNG), 2012 Ninth International Conference on
Conference_Location :
Las Vegas, NV
Print_ISBN :
978-1-4673-0798-7
Type :
conf
DOI :
10.1109/ITNG.2012.168
Filename :
6209135
Link To Document :
بازگشت