Title
Text-mining-assisted biocuration workflows in Argo.
Abstract
Biocuration activities have been broadly categorized into the selection of relevant documents, the annotation of biological concepts of interest and identification of interactions between the concepts. Text mining has been shown to have a potential to significantly reduce the effort of biocurators in all the three activities, and various semi-automatic methodologies have been integrated into curation pipelines to support them. We investigate the suitability of Argo, a workbench for building text-mining solutions with the use of a rich graphical user interface, for the process of biocuration. Central to Argo are customizable workflows that users compose by arranging available elementary analytics to form task-specific processing units. A built-in manual annotation editor is the single most used biocuration tool of the workbench, as it allows users to create annotations directly in text, as well as modify or delete annotations created by automatic processing components. Apart from syntactic and semantic analytics, the ever-growing library of components includes several data readers and consumers that support well-established as well as emerging data interchange formats such as XMI, RDF and BioC, which facilitate the inter-operability of Argo with other platforms or resources. To validate the suitability of Argo for curation activities, we participated in the BioCreative IV challenge whose purpose was to evaluate Web-based systems addressing user-defined biocuration tasks. Argo proved to have the edge over other systems in terms of flexibility of defining biocuration tasks. As expected, the versatility of the workbench inevitably lengthened the time the curators spent on learning the system before taking on the task, which may have affected the usability of Argo. The participation in the challenge gave us an opportunity to gather valuable feedback and identify areas of improvement, some of which have already been introduced. Database URL: http://argo.nactem.ac.uk
Year
DOI
Venue
2014
10.1093/database/bau070
DATABASE-THE JOURNAL OF BIOLOGICAL DATABASES AND CURATION
Keywords
Field
DocType
data mining,internet,data curation
Data mining,Interoperability,Computer science,Data curation,Analytics,Workflow,RDF,Workbench,World Wide Web,Information retrieval,Usability,Semantic analytics,Bioinformatics
Journal
Volume
ISSN
Citations 
2014
1758-0463
8
PageRank 
References 
Authors
0.46
30
5
Name
Order
Citations
PageRank
Rafal Rak138218.30
Riza Theresa Batista-Navarro29810.87
Andrew Rowley31058.87
Jacob Carter4232.50
Sophia Ananiadou52658183.08