Title
Learning to select pseudo labels: a semi-supervised method for named entity recognition
Abstract
Deep learning models have achieved state-of-the-art performance in named entity recognition (NER); the good performance, however, relies heavily on substantial amounts of labeled data. In some specific areas such as medical, financial, and military domains, labeled data is very scarce, while unlabeled data is readily available. Previous studies have used unlabeled data to enrich word representations, but a large amount of entity information in unlabeled data is neglected, which may be beneficial to the NER task. In this study, we propose a semi-supervised method for NER tasks, which learns to create high-quality labeled data by applying a pre-trained module to filter out erroneous pseudo labels. Pseudo labels are automatically generated for unlabeled data and used as if they were true labels. Our semi-supervised framework includes three steps: constructing an optimal single neural model for a specific NER task, learning a module that evaluates pseudo labels, and creating new labeled data and improving the NER model iteratively. Experimental results on two English NER tasks and one Chinese clinical NER task demonstrate that our method further improves the performance of the best single neural model. Even when we use only pre-trained static word embeddings and do not rely on any external knowledge, our method achieves comparable performance to those state-of-the-art models on the CoNLL-2003 and OntoNotes 5.0 English NER tasks.
Year
DOI
Venue
2020
10.1631/FITEE.1800743
Frontiers of Information Technology & Electronic Engineering
Keywords
DocType
Volume
Named entity recognition, Unlabeled data, Deep learning, Semi-supervised method, TP391.1
Journal
21
Issue
ISSN
Citations 
6
2095-9184
0
PageRank 
References 
Authors
0.34
0
4
Name
Order
Citations
PageRank
Zhenzhen Li124.09
Dawei Feng2112.57
Dongsheng Li3158.74
Xicheng Lu41276110.03