DocumentCode
2228822
Title
An Efficient Approach for Finding Near Duplicate Web Pages Using Minimum Weight Overlapping Method
Author
Das, Sankar Narayan ; Mathew, Michael ; Vijayaraghavan, Pramod K.
Author_Institution
Dept. of Comput. Applic., CUSAT, India
fYear
2012
fDate
16-18 April 2012
Firstpage
121
Lastpage
126
Abstract
The existence of billions of web data has severely affected the performance and reliability of web search. The presence of near duplicate web pages plays an important role in this performance degradation while integrating data from heterogeneous sources. Web mining faces huge problems due to the existence of such documents. These pages increase the index storage space and thereby increase the serving cost. By introducing efficient methods to detect and remove such documents from the Web not only decreases the computation time but also increases the relevancy of search results. We aim a novel idea for finding near duplicate web pages which can be incorporated in the field of plagiarism detection, spam detection and focused web crawling scenarios. Here we propose an efficient method for finding near duplicates of an input web page, from a huge repository. A TDW matrix based algorithm is proposed with three phases, rendering, filtering and verification, which receives an input web page and a threshold in its first phase, prefix filtering and positional filtering to reduce the size of record set in the second phase and returns an optimal set of near duplicate web pages in the verification phase by using Minimum Weight Overlapping (MWO) method. The experimental results show that our algorithm outperforms in terms of two benchmark measures, precision and recall, and a reduction in the size of competing record set.
Keywords
Internet; data mining; document handling; information filtering; rendering (computer graphics); security of data; unsolicited e-mail; TDW matrix based algorithm; Web mining; Web page finding; Web search; document detection; document removal; filtering phase; focused Web crawling scenario; minimum weight overlapping method; near duplicate Web page; plagiarism detection; positional filtering; prefix filtering; rendering phase; spam detection; verification phase; Algorithm design and analysis; Filtering; Rendering (computer graphics); Search engines; Standards; Web pages; Minimum Weight overlapping; Near Duplicate Detection; Positional filtering; Prefix filtering; Term Document Weight Matrix; web page classification;
fLanguage
English
Publisher
ieee
Conference_Titel
Information Technology: New Generations (ITNG), 2012 Ninth International Conference on
Conference_Location
Las Vegas, NV
Print_ISBN
978-1-4673-0798-7
Type
conf
DOI
10.1109/ITNG.2012.168
Filename
6209135
Link To Document