• DocumentCode
    2077513
  • Title

    A comprehensive tool for text categorization and text summarization in bioinformatics

  • Author

    Kamal, Md Mustafa ; Sultana, K.Z.

  • Author_Institution
    Dept. of Comput. Sci. & Eng., Chittagong Univ. of Eng. & Technol., Chittagong, Bangladesh
  • fYear
    2012
  • fDate
    22-24 Dec. 2012
  • Firstpage
    592
  • Lastpage
    597
  • Abstract
    The work focuses on the integration of text categorization and text summarization tasks based on some existing algorithms. We primarily employ the method for bioinformatics literatures to categorize them in relevant domains of bioinformatics and then get a summarized overview of each of the documents in the domain. For text categorization we have chosen three different and core domains of bioinformatics: Protein-Protein Interaction, Disease-Drug Relevance and Pathway-Process Involvement. The method uses TF-IDF based technology for the categorization task and then after categorization it summarizes the key contents of each document using some existing features. The system plays important role in automatically reducing review spaces for the researchers as they do not need to manually select their relevant texts. It also saves time by providing ranked and significantly relevant lines of the documents. Our method outperforms other existing summarization tools in the sense that it optimizes summarization by first categorizing the documents on the basis of TF-IDF technology and then avoids redundant information by properly ranking the sentences using existing score.
  • Keywords
    bioinformatics; diseases; drugs; proteins; text analysis; TF-IDF based technology; automatic review space reduction; bioinformatics literature; disease-drug relevance; pathway-process involvement; protein-protein interaction; redundant information; sentence ranking; summarization tool; text categorization; text summarization task; Pathway; SumBasic score; TF-IDF; Text Categorization; Text Summarization;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Computer and Information Technology (ICCIT), 2012 15th International Conference on
  • Conference_Location
    Chittagong
  • Print_ISBN
    978-1-4673-4833-1
  • Type

    conf

  • DOI
    10.1109/ICCITechn.2012.6509764
  • Filename
    6509764