• DocumentCode
    2492081
  • Title

    Host-IP Clustering Technique for Deep Web Characterization

  • Author

    Shestakov, Denis ; Salakoski, Tapio

  • Author_Institution
    Dept. of Media Technol., Helsinki Univ. of Technol., Helsinki, Finland
  • fYear
    2010
  • fDate
    6-8 April 2010
  • Firstpage
    378
  • Lastpage
    380
  • Abstract
    A huge portion of today´s Web consists of web pages filled with information from myriads of online databases. This part of the Web, known as the deep Web, is to date relatively unexplored and even major characteristics such as number of searchable databases on the Web is somewhat disputable. In this paper, we are aimed at more accurate estimation of main parameters of the deep Web by sampling one national web domain. We propose the Host-IP clustering sampling technique that addresses drawbacks of existing approaches to characterize the deep Web and report our findings based on the survey of Russian Web conducted in September 2006. Obtained estimates together with a proposed sampling method could be useful for further studies to handle data in the deep Web.
  • Keywords
    Web sites; data handling; information retrieval systems; information services; sampling methods; Russian Web; Web domain sampling; Web pages; deep Web characterization; host-IP clustering sampling; online databases; searchable databases; Databases; Internet; Parameter estimation; Protocols; Random number generation; Sampling methods; Web pages; Web search; Web server; deep web; host-IP clustering sampling; search interface discovery; virtual hosting; web characterization;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Web Conference (APWEB), 2010 12th International Asia-Pacific
  • Conference_Location
    Busan
  • Print_ISBN
    978-1-7695-4012-2
  • Electronic_ISBN
    978-1-4244-6600-9
  • Type

    conf

  • DOI
    10.1109/APWeb.2010.59
  • Filename
    5474106