DocumentCode
969238
Title
Web search engines. Part 1
Author
Hawking, David
Author_Institution
ICT Centre, CSIRO, Canberra, ACT
Volume
39
Issue
6
fYear
2006
fDate
6/1/2006 12:00:00 AM
Firstpage
86
Lastpage
88
Abstract
In this article, we go behind the scenes and explain how this data processing "miracle" is possible. We focus on whole-of-Web search but note that enterprise search tools and portal search interfaces use many of the same data structures and algorithms. Search engines cannot and should not index every page on the Web. After all, thanks to dynamic Web page generators such as automatic calendars, the number of pages is infinite. To provide a useful and cost-effective service, search engines must reject as much low-value automated content as possible. In addition, they can ignore huge volumes of Web-accessible data, such as ocean temperatures and astrophysical observations, without harm to search effectiveness. Finally, Web search engines have no access to restricted content, such as pages on corporate intranets. What follows is not an inside view of any particular commercial engine - whose precise details are jealously guarded secrets - but a characterization of the problems that whole-of-Web search services face and an explanation of the techniques available to solve these problems
Keywords
Internet; indexing; information retrieval; portals; search engines; Web page indexing; Web search engine; Web search services; data processing; data structures; dynamic Web page generator; enterprise search tool; portal search interface; Counting circuits; Crawlers; Data processing; Databases; Fingerprint recognition; Indexing; Search engines; Uniform resource locators; Web search; Web server; crawling algorithms; search engines; web crawling and searching;
fLanguage
English
Journal_Title
Computer
Publisher
ieee
ISSN
0018-9162
Type
jour
DOI
10.1109/MC.2006.213
Filename
1642621
Link To Document