Title
Scholarly big data information extraction and integration in the CiteSeerχ digital library
Abstract
CiteSeerχ is a digital library that contains approximately 3.5 million scholarly documents and receives between 2 and 4 million requests per day. In addition to making documents available via a public Website, the data is also used to facilitate research in areas like citation analysis, co-author network analysis, scalability evaluation and information extraction. The papers in CiteSeerχ are gathered from the Web by means of continuous automatic focused crawling and go through a series of automatic processing steps as part of the ingestion process. Given the size of the collection, the fact that it is constantly expanding, and the multiple ways in which it is used both by the public to access scholarly documents and for research, there are several big data challenges. In this paper, we provide a case study description of how we address these challenges when it comes to information extraction, data integration and entity linking in CiteSeerχ. We describe how we: aggregate data from multiple sources on the Web; store and manage data; process data as part of an automatic ingestion pipeline that includes automatic metadata and information extraction; perform document and citation clustering; perform entity linking and name disambiguation; and make our data and source code available to enable research and collaboration.
Year
DOI
Venue
2014
10.1109/ICDEW.2014.6818305
Data Engineering Workshops
Keywords
DocType
Citations 
Big Data,Web sites,citation analysis,data integration,digital libraries,information retrieval,meta data,pattern clustering,CiteSeerχ digital library,automatic ingestion pipeline,automatic metadata,automatic processing steps,big data information extraction,big data information integration,citation analysis,information extraction,public Website
Conference
0
PageRank 
References 
Authors
0.34
0
5
Name
Order
Citations
PageRank
Kyle Williams120821.61
Jian Wu2226.11
Sagnik Ray Choudhury3745.86
Madian Khabsa423718.81
C. Lee Giles5111541549.48