• Title of article

    A data labeling method for clustering categorical data

  • Author/Authors

    Cao، نويسنده , , Fuyuan and Liang، نويسنده , , Jiye، نويسنده ,

  • Issue Information
    روزنامه با شماره پیاپی سال 2011
  • Pages
    5
  • From page
    2381
  • To page
    2385
  • Abstract
    As the size of data growing at a rapid pace, clustering a very large data set inevitably incurs a time-consuming process. To improve the efficiency of clustering, sampling is usually used to scale down the size of data set. However, with sampling applied, how to allocate unlabeled objects into proper clusters is a very difficult problem. In this paper, based on the frequency of attribute values in a given cluster and the distributions of attribute values in different clusters, a novel similarity measure is proposed to allocate each unlabeled object into the corresponding appropriate cluster for clustering categorical data. Furthermore, a labeling algorithm for categorical data is presented, and its corresponding time complexity is analyzed as well. The effectiveness of the proposed algorithm is shown by the experiments on real-world data sets.
  • Keywords
    Data labeling , Rough membership function , Similarity measure , Categorical data
  • Journal title
    Expert Systems with Applications
  • Serial Year
    2011
  • Journal title
    Expert Systems with Applications
  • Record number

    2348882