• DocumentCode
    2228822
  • Title

    An Efficient Approach for Finding Near Duplicate Web Pages Using Minimum Weight Overlapping Method

  • Author

    Das, Sankar Narayan ; Mathew, Michael ; Vijayaraghavan, Pramod K.

  • Author_Institution
    Dept. of Comput. Applic., CUSAT, India
  • fYear
    2012
  • fDate
    16-18 April 2012
  • Firstpage
    121
  • Lastpage
    126
  • Abstract
    The existence of billions of web data has severely affected the performance and reliability of web search. The presence of near duplicate web pages plays an important role in this performance degradation while integrating data from heterogeneous sources. Web mining faces huge problems due to the existence of such documents. These pages increase the index storage space and thereby increase the serving cost. By introducing efficient methods to detect and remove such documents from the Web not only decreases the computation time but also increases the relevancy of search results. We aim a novel idea for finding near duplicate web pages which can be incorporated in the field of plagiarism detection, spam detection and focused web crawling scenarios. Here we propose an efficient method for finding near duplicates of an input web page, from a huge repository. A TDW matrix based algorithm is proposed with three phases, rendering, filtering and verification, which receives an input web page and a threshold in its first phase, prefix filtering and positional filtering to reduce the size of record set in the second phase and returns an optimal set of near duplicate web pages in the verification phase by using Minimum Weight Overlapping (MWO) method. The experimental results show that our algorithm outperforms in terms of two benchmark measures, precision and recall, and a reduction in the size of competing record set.
  • Keywords
    Internet; data mining; document handling; information filtering; rendering (computer graphics); security of data; unsolicited e-mail; TDW matrix based algorithm; Web mining; Web page finding; Web search; document detection; document removal; filtering phase; focused Web crawling scenario; minimum weight overlapping method; near duplicate Web page; plagiarism detection; positional filtering; prefix filtering; rendering phase; spam detection; verification phase; Algorithm design and analysis; Filtering; Rendering (computer graphics); Search engines; Standards; Web pages; Minimum Weight overlapping; Near Duplicate Detection; Positional filtering; Prefix filtering; Term Document Weight Matrix; web page classification;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Information Technology: New Generations (ITNG), 2012 Ninth International Conference on
  • Conference_Location
    Las Vegas, NV
  • Print_ISBN
    978-1-4673-0798-7
  • Type

    conf

  • DOI
    10.1109/ITNG.2012.168
  • Filename
    6209135