مرکز منطقه ای اطلاع رساني علوم و فناوري - A Novel and Efficient Approach For Near Duplicate Page Detection in Web Crawling

DocumentCode :

3076801

Title :

A Novel and Efficient Approach For Near Duplicate Page Detection in Web Crawling

Author :

Narayana, V.A. ; Premchand, P. ; Govardhan, A.

Author_Institution :

CSE Dept., JNTU, Hyderabad

fYear :

2009

fDate :

6-7 March 2009

Firstpage :

1492

Lastpage :

1496

Abstract :

The drastic development of the World Wide Web in the recent times has made the concept of Web crawling receive remarkable significance. The voluminous amounts of Web documents swarming the Web have posed huge challenges to the Web search engines making their results less relevant to the users. The presence of duplicate and near duplicate Web documents in abundance has created additional overheads for the search engines critically affecting their performance and quality. The detection of duplicate and near duplicate Web pages has long been recognized in Web crawling research community. It is an important requirement for search engines to provide users with the relevant results for their queries in the first page without duplicate and redundant results. In this paper, we have presented a novel and efficient approach for the detection of near duplicate Web pages in Web crawling. Detection of near duplicate Web pages is carried out ahead of storing the crawled Web pages in to repositories. At first, the keywords are extracted from the crawled pages and the similarity score between two pages is calculated based on the extracted keywords. The documents having similarity scores greater than a threshold value are considered as near duplicates. The detection has resulted in reduced memory for repositories and improved search engine quality.

Keywords :

Internet; data mining; query processing; search engines; Web crawling; Web document swarming; Web mining; Web query; Web search engine; World Wide Web; keyword extraction; near duplicate page detection; Computer science; Crawlers; Data mining; Educational institutions; Focusing; Search engines; Tellurium; Web pages; Web search; Web sites; Common words; Near duplicate detection; Near duplicate pages; Stemming; Web Content Mining; Web Crawling; Web Mining; Web pages;

fLanguage :

English

Publisher :

ieee

Conference_Titel :

Advance Computing Conference, 2009. IACC 2009. IEEE International

Conference_Location :

Patiala

Print_ISBN :

978-1-4244-2927-1

Electronic_ISBN :

978-1-4244-2928-8

Type :

conf

DOI :

10.1109/IADCC.2009.4809238

Filename :

4809238

Link To Document :

https://search.ricest.ac.ir/dl/search/defaultta.aspx?DTC=49&DC=3076801