DocumentCode
238150
Title
Design and implementation of competent web crawler and indexer using web services
Author
Santhosh, Kumar D. K. ; Kamath, Manjunath
Author_Institution
Dept. of Comput. Sci. & Eng., NMAM Inst. of Technol., Nitte, India
fYear
2014
fDate
8-10 May 2014
Firstpage
1672
Lastpage
1677
Abstract
Today the internet has become a part of human beings life. To get the information what the user is requesting is the job of search engine which indeed takes the help of web crawler. Designing and developing a competent web crawler is a challenging task. This paper proposes Web crawler and Indexer. The WebCrawler consist of crawler services and indexer services and realized as web services. The crawler and indexer services communicate using XML, SOAP and WSDL. The web pages are fetched and parsed for retrieving all the hyperlinks by the crawler service, and then the same process is continued recursively using the Breadth-First strategy. The result of crawler service is downloaded and given as an input to the indexer services by passing the message using web services. Then the indexer service parses the HTML pages, removes stop words, stemming of keywords are carried out as pre-processing steps. Finally the result is stored in the form of inverted index.
Keywords
Web services; XML; indexing; information retrieval; search engines; HTML pages; Internet; SOAP; WSDL; Web crawler; Web indexer; Web services; XML; breadth-first strategy; hyperlink retrieval; inverted index; search engine; Crawlers; HTML; Search engines; Simple object access protocol; Uniform resource locators; Web pages; Breadth first Strategy; Tokenization; hyperlink; stemming; stop-words;
fLanguage
English
Publisher
ieee
Conference_Titel
Advanced Communication Control and Computing Technologies (ICACCCT), 2014 International Conference on
Conference_Location
Ramanathapuram
Print_ISBN
978-1-4799-3913-8
Type
conf
DOI
10.1109/ICACCCT.2014.7019393
Filename
7019393
Link To Document