Title of article
An efficient and scalable family of algorithms for combining clusterings
Author/Authors
Mimaroglu، نويسنده , , Selim and Erdil، نويسنده , , Ertunc، نويسنده ,
Issue Information
روزنامه با شماره پیاپی سال 2013
Pages
15
From page
2525
To page
2539
Abstract
Clustering is the process of grouping objects that are similar, where similarity between objects is usually measured by a distance metric. The groups formed by a clustering method are referred as clusters. Clustering is a widely used activity with multiple applications ranging from biology to economics. Each clustering technique has some advantages and disadvantages. Some clustering algorithms may even require input parameters which strongly affect the result. In most cases, it is not possible to choose the best distance metric, the best clustering method, and the best input argument values for an input data set. Therefore, multiple clusterings can be obtained by several distance metrics, several clustering methods, and several input argument values. And, multiple clusterings can be combined into a new and better quality final clustering. We propose a family of combining multiple clustering algorithms that are memory efficient, scalable, robust, and intuitive. Our new algorithms offer tremendous speed gain and low memory requirements by working at cluster level, while producing very good quality final clusters. Extensive experimental evaluations on some very challenging artificially generated and real data sets from a diverse set of domains establish the usefulness of our methods.
Keywords
Fast clustering , Clustering , Combining multiple clusterings , Cluster fusion , Cluster ensemble
Journal title
Engineering Applications of Artificial Intelligence
Serial Year
2013
Journal title
Engineering Applications of Artificial Intelligence
Record number
2126046
Link To Document