Title
Exploring Attribute Correspondences Across Heterogeneous Databases by Mutual Information
Abstract
Identifying attribute correspondences across heterogeneous databases is a critical and time-consuming step in integrating the databases. Past research has applied correlation analysis techniques to explore correspondences between attributes. These techniques, however, are appropriate for numeric attributes that are linearly related. This paper proposes an information-theoretic approach to exploring correspondences between attributes in heterogeneous databases. The proposed approach is applicable to character attributes, as well as to numeric attributes, regardless whether or not they are linearly related. It overcomes some serious shortcomings of previous approaches based on correlation analysis and has much broader applicability. The proposed procedure samples both matching and nonmatching pairs of records from the databases under consideration, applies matching functions to compare pairs of attributes, and then uses the mutual information to measure the dependency between a matching function as applied to a pair of attributes and the class (i.e., matching or nonmatching) of a pair of records. A high mutual information index implies a potential attribute correspondence, which is presented to the analyst for further evaluation. The paper also presents some empirical results demonstrating the utility of the proposed approach.
Year
DOI
Venue
2006
10.2753/MIS0742-1222220411
J. of Management Information Systems
Keywords
DocType
Volume
Heterogeneous Databases,Mutual Information,correlation analysis technique,proposed approach,correlation analysis,Exploring Attribute Correspondences,Identifying attribute correspondence,character attribute,proposed procedure sample,heterogeneous databases,matching function,information-theoretic approach,previous approach
Journal
22
Issue
ISSN
Citations 
4
0742-1222
2
PageRank 
References 
Authors
0.37
42
2
Name
Order
Citations
PageRank
Huimin Zhao11108.18
Ehsan S. Soofi2678.95