Title
CALBC silver standard corpus.
Abstract
The CALBC initiative aims to provide a large-scale biomedical text corpus that contains semantic annotations for named entities of different kinds. The generation of this corpus requires that the annotations from different automatic annotation systems be harmonized. In the first phase, the annotation systems from five participants (EMBL-EBI, EMC Rotterdam, NLM, JULIE Lab Jena, and Linguamatics) were gathered. All annotations were delivered in a common annotation format that included concept identifiers in the boundary assignments and that enabled comparison and alignment of the results. During the harmonization phase, the results produced from those different systems were integrated in a single harmonized corpus ("silver standard" corpus) by applying a voting scheme. We give an overview of the processed data and the principles of harmonization--formal boundary reconciliation and semantic matching of named entities. Finally, all submissions of the participants were evaluated against that silver standard corpus. We found that species and disease annotations are better standardized amongst the partners than the annotations of genes and proteins. The raw corpus is now available for additional named entity annotations. Parts of it will be made available later on for a public challenge. We expect that we can improve corpus building activities both in terms of the numbers of named entity classes being covered, as well as the size of the corpus in terms of annotated documents.
Year
DOI
Venue
2010
10.1142/S0219720010004562
J. Bioinformatics and Computational Biology
Keywords
Field
DocType
biomedical text mining
Annotation,Harmonization,Identifier,Information retrieval,Computer science,Text corpus,Biomedical text mining,Natural language processing,Artificial intelligence,Named-entity recognition,Unified Medical Language System,Semantic matching
Journal
Volume
Issue
ISSN
8
1
1757-6334
Citations 
PageRank 
References 
31
1.36
8
Authors
10
Name
Order
Citations
PageRank
dietrich rebholzschuhmann1102375.06
Antonio Jimeno-Yepes254033.38
Erik M. Van Mulligen363344.63
Ning Kang4704.36
Jan Kors5815.64
David Milward619627.51
Peter Corbett717613.18
Ekaterina Buyko816511.45
Elena Beisswanger922316.00
Udo Hahn1093788.14