• DocumentCode
    2727967
  • Title

    Concept Forest: A New Ontology-assisted Text Document Similarity Measurement Method

  • Author

    Wang, James Z. ; Taylor, William

  • fYear
    2007
  • fDate
    2-5 Nov. 2007
  • Firstpage
    395
  • Lastpage
    401
  • Abstract
    Although using ontologies to assist information retrieval and text document processing has recently attracted more and more attention, existing ontologybased approaches have not shown advantages over the traditional keywords-based Latent Semantic Indexing (LSI) method. This paper proposes an algorithm to extract a concept forest (CF) from a document with the assistance of a natural language ontology, the WordNet lexical database. Using concept forests to represent the semantics of text documents, the semantic similarities of these documents are then measured as the commonalities of their concept forests. Performance studies of text document clustering based on different document similarity measurement methods show that the CF-based similarity measurement is an effective alternative to the existing keywords-based methods. In particular, this CFbased approach has obvious advantages over the existing keywords-based methods, including LSI, in processing short text documents or in P2P or live news environments where it is impractical to collect the entire document corpus for analysis.
  • Keywords
    Feeds; Frequency; Indexing; Information retrieval; Large scale integration; Natural languages; Ontologies; Text analysis; Text mining; USA Councils;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Web Intelligence, IEEE/WIC/ACM International Conference on
  • Conference_Location
    Fremont, CA
  • Print_ISBN
    978-0-7695-3026-0
  • Type

    conf

  • DOI
    10.1109/WI.2007.11
  • Filename
    4427122