Abstract | ||
---|---|---|
Mining arbitrary shaped clusters in large data sets is an open challenge in data mining. Various approaches to this problem have been proposed with high time complexity. To save computational cost, some algorithms try to shrink a data set size to a smaller amount of representative data examples. However, their user-defined shrinking ratios may significantly affect the clustering performance. In this paper, we present CLASP an effective and efficient algorithm for mining arbitrary shaped clusters. It automatically shrinks the size of a data set while effectively preserving the shape information of clusters in the data set with representative data examples. Then, it adjusts the positions of these representative data examples to enhance their intrinsic relationship and make the cluster structures more clear and distinct for clustering. Finally, it performs agglomerative clustering to identify the cluster structures with the help of a mutual k-nearest neighbors-based similarity metric called Pk. Extensive experiments on both synthetic and real data sets are conducted, and the results verify the effectiveness and efficiency of our approach. |
Year | DOI | Venue |
---|---|---|
2014 | 10.1109/ICDE.2014.6816637 | ICDE |
Keywords | Field | DocType |
clustering performance,pattern clustering,k-nearest neighbors-based similarity metric,user-defined shrinking ratios,time complexity,real data sets,computational cost,agglomerative clustering,arbitrary shaped clusters mining,clasp,shape information,large data sets,data set size,data mining,cluster structures,synthetic data sets,shape,algorithm design and analysis,symmetric matrices,clustering algorithms | k-medians clustering,Data mining,CURE data clustering algorithm,Affinity propagation,Correlation clustering,Computer science,Determining the number of clusters in a data set,Constrained clustering,Cluster analysis,Single-linkage clustering | Conference |
ISSN | Citations | PageRank |
1084-4627 | 8 | 0.52 |
References | Authors | |
16 | 5 |
Name | Order | Citations | PageRank |
---|---|---|---|
Hao Huang | 1 | 89 | 7.77 |
Yunjun Gao | 2 | 862 | 89.71 |
Kevin Chiew | 3 | 116 | 11.06 |
Lei Chen | 4 | 6239 | 395.84 |
Qinming He | 5 | 371 | 41.53 |