• Title of article

    A logistic regression-based smoothing method for Chinese text categorization

  • Author/Authors

    Yen، نويسنده , , Show-Jane and Lee، نويسنده , , Yue-Shi and Ying، نويسنده , , Jia-Ching and Wu، نويسنده , , Yu-Chieh، نويسنده ,

  • Issue Information
    روزنامه با شماره پیاپی سال 2011
  • Pages
    10
  • From page
    11581
  • To page
    11590
  • Abstract
    Automatic Chinese text classification is an important and a well-known technology in the field of machine learning. The first step for solving Chinese text categorization problems is to tokenize the Chinese words from a sequence of non-segmented sentences. However, previous literatures often employ a Chinese word tokenizer that was trained with different sources and then perform the conventional text classification approaches. However, these taggers are not perfect and often provide incorrect word boundary information. In this paper, we propose an N-gram-based language model which takes word relations into account for Chinese text categorization without Chinese word tokenizer. To prevent from out-of-vocabulary, we also propose a novel smoothing approach based on logistic regression to improve accuracy. The experimental result shows that our approach outperforms traditional methods at least 11% on micro-average F-measure.
  • Keywords
    Text classification , N-gram-based classification , feature selection , Word segmentation , logistic regression
  • Journal title
    Expert Systems with Applications
  • Serial Year
    2011
  • Journal title
    Expert Systems with Applications
  • Record number

    2350099