• DocumentCode
    2711204
  • Title

    Using Wikipedia for Co-clustering Based Cross-Domain Text Classification

  • Author

    Wang, Pu ; Domeniconi, Carlotta ; Hu, Jian

  • Author_Institution
    Dept. of Comput. Sci., George Mason Univ., Fairfax, VA
  • fYear
    2008
  • fDate
    15-19 Dec. 2008
  • Firstpage
    1085
  • Lastpage
    1090
  • Abstract
    Traditional approaches to document classification requires labeled data in order to construct reliable and accurate classifiers. Unfortunately, labeled data are seldom available, and often too expensive to obtain. Given a learning task for which training data are not available, abundant labeled data may exist for a different but related domain. One would like to use the related labeled data as auxiliary information to accomplish the classification task in the target domain. Recently, the paradigm of transfer learning has been introduced to enable effective learning strategies when auxiliary data obey a different probability distribution. A co-clustering based classification algorithm has been previously proposed to tackle cross-domain text classification. In this work, we extend the idea underlying this approach by making the latent semantic relationship between the two domains explicit. This goal is achieved with the use of Wikipedia. As a result, the pathway that allows to propagate labels between the two domains not only captures common words, but also semantic concepts based on the content of documents. We empirically demonstrate the efficacy of our semantic-based approach to cross-domain classification using a variety of real data.
  • Keywords
    Web sites; classification; learning (artificial intelligence); pattern clustering; statistical distributions; text analysis; co-clustering-based cross-domain text classification; document classification; latent semantic relationship; learning strategy; probability distribution; wikipedia; Asia; Bridges; Classification algorithms; Computer science; Data mining; Dictionaries; Probability distribution; Text categorization; Training data; Wikipedia; Co-clustering; Cross-domain text classification; Transfer learning; Wikipedia;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Data Mining, 2008. ICDM '08. Eighth IEEE International Conference on
  • Conference_Location
    Pisa
  • ISSN
    1550-4786
  • Print_ISBN
    978-0-7695-3502-9
  • Type

    conf

  • DOI
    10.1109/ICDM.2008.136
  • Filename
    4781229