Title
Kurator: A Kepler Package For Data Curation Workflows
Abstract
Data curation is critical for scientific data digitization, sharing, integration, and use. This paper presents Kurator, a software package for automating data curation pipelines in the Kepler scientific workflow system. Several curation tools and services are integrated into this package as actors to enable construction of workflows to perform and document various data curation tasks. The integration of Google cloud services (e. g., Google spreadsheets), allows workflow steps to invoke human experts outside the workflow in a manner that greatly simplifies the complex data handling in distributed, multi-user curation workflows. The Kepler platform provides the modeling, execution and management ability, including a collection-oriented model of computation (COMAD), and provenance tracking and browsing for the curation package. These features not only allow workflows to be easily modeled, maintained, and evolved, but also QA/QC of curation results is facilitated through examination of provenance information recorded during workflow execution. Effectiveness of the Kurator package is demonstrated through a workflow for data curation of natural science collections.
Year
DOI
Venue
2012
10.1016/j.procs.2012.04.177
PROCEEDINGS OF THE INTERNATIONAL CONFERENCE ON COMPUTATIONAL SCIENCE, ICCS 2012
Keywords
Field
DocType
data curation, scientific workflows, biodiversity informatics
Data mining,Digitization,World Wide Web,Biodiversity informatics,Kepler scientific workflow system,Software engineering,Computer science,Data curation,Software,Model of computation,Workflow,Cloud computing
Journal
Volume
ISSN
Citations 
9
1877-0509
4
PageRank 
References 
Authors
0.50
8
7
Name
Order
Citations
PageRank
L. Dou140.50
G. Cao240.50
Paul J. Morris391.34
Robert A. Morris440.50
B. LudSscher540.50
James A. Macklin6182.82
J. Hanken7292.68