Title
Automatically incorporating new sources in keyword search-based data integration
Abstract
Scientific data offers some of the most interesting challenges in data integration today. Scientific fields evolve rapidly and accumulate masses of observational and experimental data that needs to be annotated, revised, interlinked, and made available to other scientists. From the perspective of the user, this can be a major headache as the data they seek may initially be spread across many databases in need of integration. Worse, even if users are given a solution that integrates the current state of the source databases, new data sources appear with new data items of interest to the user. Here we build upon recent ideas for creating integrated views over data sources using keyword search techniques, ranked answers, and user feedback [32] to investigate how to automatically discover when a new data source has content relevant to a user's view - in essence, performing automatic data integration for incoming data sets. The new architecture accommodates a variety of methods to discover related attributes, including label propagation algorithms from the machine learning community [2] and existing schema matchers [11]. The user may provide feedback on the suggested new results, helping the system repair any bad alignments or increase the cost of including a new source that is not useful. We evaluate our approach on actual bioinformatics schemas and data, using state-of-the-art schema matchers as components. We also discuss how our architecture can be adapted to more traditional settings with a mediated schema.
Year
DOI
Venue
2010
10.1145/1807167.1807211
SIGMOD Conference
Keywords
Field
DocType
keyword search-based data integration,new result,incoming data set,automatic data integration,new architecture,data integration,new data source,scientific data,new source,new data item,experimental data,data source,machine learning,data integrity
Data integration,Data warehouse,Data mining,Data architecture,Data exchange,Experimental data,Ranking,Computer science,Schema matching,Schema (psychology),Database
Conference
Citations 
PageRank 
References 
20
0.82
35
Authors
3
Name
Order
Citations
PageRank
Partha Pratim Talukdar198065.47
Zachary G. Ives23869318.82
Fernando Pereira3177172124.79