Title
Hierarchical document clustering using local patterns
Abstract
The global pattern mining step in existing pattern-based hierarchical clustering algorithms may result in an unpredictable number of patterns. In this paper, we propose IDHC, a pattern-based hierarchical clustering algorithm that builds a cluster hierarchy without mining for globally significant patterns. IDHC first discovers locally promising patterns by allowing each instance to "vote" for its representative size-2 patterns in a way that ensures an effective balance between local pattern frequency and pattern significance in the dataset. The cluster hierarchy (i.e., the global model) is then directly constructed using these locally promising patterns as features. Each pattern forms an initial (possibly overlapping) cluster, and the rest of the cluster hierarchy is obtained by following a unique iterative cluster refinement process. By effectively utilizing instance-to-cluster relationships, this process directly identifies clusters for each level in the hierarchy, and efficiently prunes duplicate clusters. Furthermore, IDHC produces cluster labels that are more descriptive (patterns are not artificially restricted), and adapts a soft clustering scheme that allows instances to exist in suitable nodes at various levels in the cluster hierarchy. We present results of experiments performed on 16 standard text datasets, and show that IDHC outperforms state-of-the-art hierarchical clustering algorithms in terms of average entropy and FScore measures.
Year
DOI
Venue
2010
10.1007/s10618-010-0172-z
Data Min. Knowl. Discov.
Keywords
DocType
Volume
Pattern based hierarchical clustering,Interestingness measures,Dimensionality reduction,Pattern selection,Global modeling using local patterns
Journal
21
Issue
ISSN
Citations 
1
1384-5810
15
PageRank 
References 
Authors
0.59
20
4
Name
Order
Citations
PageRank
Hassan H. Malik1775.10
John R. Kender2627138.04
Dmitriy Fradkin334419.25
Fabian Moerchen417210.51