• DocumentCode
    1688232
  • Title

    Linear and sublinear time algorithms for mining frequent traversal path patterns from very large Web logs

  • Author

    Chen, Zhixiang ; Fowler, Richard H. ; Fu, Ada Wai-Chee ; Wang, Chunyue

  • Author_Institution
    Dept. of Comput. Sci., Univ. of Texas, Edinburg, TX, USA
  • fYear
    2003
  • Firstpage
    117
  • Lastpage
    122
  • Abstract
    This paper aims for designing algorithms for the problem of mining frequent traversal path patterns from very large Web logs with best possible efficiency. We devise two algorithms for this problem with the help of fast construction of "shallow" generalized suffix trees over a very large alphabet. These two algorithms have respectively provable linear time and sublinear complexity, and their performance is analyzed in comparison with the two a priori-like algorithms in (Chen et al., 1998) and the well-known Ukkonen algorithm for online suffix tree construction (1995). It is shown that these two algorithms are substantially efficient than the two apriori-like algorithms and the Ukkonen algorithm. The linear time algorithm has optimal performance in theory, while the sublinear time algorithm has better empirical performance.
  • Keywords
    Internet; computational complexity; data mining; database management systems; pattern recognition; program verification; tree data structures; Ukkonen algorithm; empirical performance; frequent traversal path pattern mining; graph structure; large alphabet; linear time algorithm; online suffix tree construction; optimal performance; performance analysis; shallow generalized suffix tree; sublinear time; time complexity; very large Web logs; Algorithm design and analysis; Computer science; Data engineering; Databases; Frequency; Performance analysis; Tree graphs;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Database Engineering and Applications Symposium, 2003. Proceedings. Seventh International
  • ISSN
    1098-8068
  • Print_ISBN
    0-7695-1981-4
  • Type

    conf

  • DOI
    10.1109/IDEAS.2003.1214918
  • Filename
    1214918