• DocumentCode
    3108402
  • Title

    TFIDF, LSI and multi-word in information retrieval and text categorization

  • Author

    Zhang, Wen ; Yoshida, Taketoshi ; Tang, Xijin

  • Author_Institution
    Sch. of Knowledge Sci., Japan Adv. Inst. of Sci. & Technol., Tatsunokuchi
  • fYear
    2008
  • fDate
    12-15 Oct. 2008
  • Firstpage
    108
  • Lastpage
    113
  • Abstract
    Text representation, which is a fundamental and necessary process for text-based intelligent information processing, includes the tasks of determining the index terms for documents and producing the numeric vectors corresponding to the documents. In this paper, multi-word, which is regarded as containing more contextual semantics than individual word and possessing the favorable statistical characteristics, is proposed as an alternative index terms in vector space model for text representation with theoretical support. We investigate the traditional indexing methods as TF*IDF (term frequency inverse document frequency) and LSI (latent semantic indexing) for comparative study. The performances of TF*IDF, LSI and multi-word are examined on the tasks of text classification, which includes information retrieval (IR) and text categorization (TC), in Chinese and English document collection respectively. We also attempt to tune the rescaling factor of LSI and observe its effectiveness in text classification. The experimental results demonstrate that TF*IDF and multi-word are comparable when they are used for IR and TC and LSI is the poorest one of them. Moreover, the rescaling factor of LSI has an insignificant influence on its effectiveness on text classification for both Chinese and English text classification.
  • Keywords
    information retrieval; text analysis; LSI; TFIDF; contextual semantics; information retrieval; latent semantic indexing; rescaling factor; term frequency inverse document frequency; text categorization; text classification; text representation; text-based intelligent information processing; traditional indexing methods; vector space model; Context modeling; Data mining; Frequency; Indexing; Information processing; Information retrieval; Large scale integration; Mathematics; Text categorization; Text mining; LSI; TF*IDF; multi-word; text classification; text representation;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Systems, Man and Cybernetics, 2008. SMC 2008. IEEE International Conference on
  • Conference_Location
    Singapore
  • ISSN
    1062-922X
  • Print_ISBN
    978-1-4244-2383-5
  • Electronic_ISBN
    1062-922X
  • Type

    conf

  • DOI
    10.1109/ICSMC.2008.4811259
  • Filename
    4811259