The Newspaper Navigator Dataset: Extracting Headlines and Visual Content from 16 Million Historic Newspaper Pages in Chronicling America - Citegraph

Paper Info

Title
The Newspaper Navigator Dataset: Extracting Headlines and Visual Content from 16 Million Historic Newspaper Pages in Chronicling America

Abstract
Chronicling America is a product of the National Digital Newspaper Program, a partnership between the Library of Congress and the National Endowment for the Humanities to digitize historic American newspapers. Over 16 million pages have been digitized to date, complete with high-resolution images and machine-readable METS/ALTO OCR. Of considerable interest to Chronicling America users is a semantified corpus, complete with extracted visual content and headlines. To accomplish this, we introduce a visual content recognition model trained on bounding box annotations collected as part of the Library of Congress's Beyond Words crowdsourcing initiative and augmented with additional annotations including those of headlines and advertisements. We describe our pipeline that utilizes this deep learning model to extract 7 classes of visual content: headlines, photographs, illustrations, maps, comics, editorial cartoons, and advertisements, complete with textual content such as captions derived from the METS/ALTO OCR, as well as image embeddings. We report the results of running the pipeline on 16.3 million pages from the Chronicling America corpus and describe the resulting Newspaper Navigator dataset, the largest dataset of extracted visual content from historic newspapers ever produced. The Newspaper Navigator dataset, finetuned visual content recognition model, and all source code are placed in the public domain for unrestricted re-use.

Year	DOI	Venue
2020	10.1145/3340531.3412767	arxiv
DocType	ISBN	Citations
Conference	978-1-4503-6859-9	0
PageRank	References	Authors
0.34	0	10

Authors (10 rows)

Cited by (0 rows)

References (0 rows)

Name	Order	Citations	PageRank
Germain Lee Benjamin Charles	1	0	0.34
Jaime Mears	2	0	0.34
Eileen Jakeway	3	0	0.34
Meghan Ferriter	4	0	0.34
Chris Adams	5	0	0.34
Nathan Yarasavage	6	0	0.68
Deborah Thomas	7	0	0.34
Kate Zwaard	8	0	0.34
Daniel S. Weld	9	10298	1127.49
Benjamin Charles Germain Lee	10	0	0.34

1