Title
Deep Visual Template-Free Form Parsing
Abstract
Automatic, template-free extraction of information from form images is challenging due to the variety of form layouts. This is even more challenging for historical forms due to noise and degradation. A crucial part of the extraction process is associating input text with pre-printed labels. We present a learned, template-free solution to detecting pre-printed text and input text/handwriting and predicting pair-wise relationships between them. While previous approaches to this problem have been focused on clean images and clear layouts, we show our approach is effective in the domain of noisy, degraded, and varied form images. We introduce a new dataset of historical form images (late 1800s, early 1900s) for training and validating our approach. Our method uses a convolutional network to detect pre-printed text and input text lines. We pool features from the detection network to classify possible relationships in a language-agnostic way. We show that our proposed pairing method outperforms heuristic rules and that visual features are critical to obtaining high accuracy.
Year
DOI
Venue
2019
10.1109/ICDAR.2019.00030
2019 International Conference on Document Analysis and Recognition (ICDAR)
Keywords
Field
DocType
template-free,document understanding,forms,form understanding,form parsing,pairing,historical
Heuristic,Pattern recognition,Handwriting,Computer science,Pairing,Artificial intelligence,Parsing,Free form
Conference
ISSN
ISBN
Citations 
1520-5363
978-1-7281-3015-6
0
PageRank 
References 
Authors
0.34
0
5
Name
Order
Citations
PageRank
Brian Davis16612.87
Bryan S. Morse267290.28
Scott Cohen3105649.07
Brian Price470933.13
Chris Tensmeyer5204.83