• DocumentCode
    2769098
  • Title

    Predicting Web Search Hit Counts

  • Author

    Tian, Tian ; Geller, James ; Chun, Soon Ae

  • Author_Institution
    New Jersey Inst. of Technol., Newark, NJ, USA
  • Volume
    1
  • fYear
    2010
  • fDate
    Aug. 31 2010-Sept. 3 2010
  • Firstpage
    162
  • Lastpage
    166
  • Abstract
    Keyword-based search engines often return an unexpected number of results. Zero hits are naturally undesirable, while too many hits are likely to be overwhelming and of low precision. We present an approach for predicting the number of hits for a given set of query terms. Using word frequencies derived from a large corpus, we construct random samples of combinations of these words as search terms. Then we derive a correlation function between the computed probabilities of search terms and the observed hit counts for them. This regression function is used to predict the hit counts for a user´s new searches, with the intention of avoiding information overload. We report the results of experiments with Google, Yahoo! and Bing to validate our methodology. We further investigate the monotonicity of search results for negative search terms by those three search engines.
  • Keywords
    Internet; content-based retrieval; information retrieval; regression analysis; search engines; Bing; Google; Web search hit count; Yahoo; keyword-based search engine; query term; regression function;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Web Intelligence and Intelligent Agent Technology (WI-IAT), 2010 IEEE/WIC/ACM International Conference on
  • Conference_Location
    Toronto, ON
  • Print_ISBN
    978-1-4244-8482-9
  • Electronic_ISBN
    978-0-7695-4191-4
  • Type

    conf

  • DOI
    10.1109/WI-IAT.2010.227
  • Filename
    5616242