• DocumentCode
    3122748
  • Title

    GraphSig: A Scalable Approach to Mining Significant Subgraphs in Large Graph Databases

  • Author

    Ranu, Sayan ; Singh, Ambuj K.

  • Author_Institution
    Dept. of Comput. Sci., Univ. of California, Santa Barbara, CA
  • fYear
    2009
  • fDate
    March 29 2009-April 2 2009
  • Firstpage
    844
  • Lastpage
    855
  • Abstract
    Graphs are being increasingly used to model a wide range of scientific data. Such widespread usage of graphs has generated considerable interest in mining patterns from graph databases. While an array of techniques exists to mine frequent patterns, we still lack a scalable approach to mine statistically significant patterns, specifically patterns with low p-values, that occur at low frequencies. We propose a highly scalable technique, called GraphSig, to mine significant subgraphs from large graph databases. We convert each graph into a set of feature vectors where each vector represents a region within the graph. Domain knowledge is used to select a meaningful feature set. Prior probabilities of features are computed empirically to evaluate statistical significance of patterns in the feature space. Following analysis in the feature space, only a small portion of the exponential search space is accessed for further analysis. This enables the use of existing frequent subgraph mining techniques to mine significant patterns in a scalable manner even when they are infrequent. Extensive experiments are carried out on the proposed techniques, and empirical results demonstrate that GraphSig is effective and efficient for mining significant patterns. To further demonstrate the power of significant patterns, we develop a classifier using patterns mined by GraphSig. Experimental results show that the proposed classifier achieves superior performance, both in terms of quality and computation cost, over state-of-the-art classifiers.
  • Keywords
    data mining; graph theory; very large databases; GraphSig; data mining; domain knowledge; large graph databases; Chemical compounds; Chemical technology; Computational efficiency; Computer science; Data engineering; Frequency; Probability; Social network services; Spatial databases; USA Councils; Chemical Compound Classification; Graph Mining; Significant Subgraphs;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Data Engineering, 2009. ICDE '09. IEEE 25th International Conference on
  • Conference_Location
    Shanghai
  • ISSN
    1084-4627
  • Print_ISBN
    978-1-4244-3422-0
  • Electronic_ISBN
    1084-4627
  • Type

    conf

  • DOI
    10.1109/ICDE.2009.133
  • Filename
    4812459