Title
Internet Argument Corpus 2.0: An SQL schema for Dialogic Social Media and the Corpora to go with it.
Abstract
Large scale corpora have benefited many areas of research in natural language processing, but until recently, resources for dialogue have lagged behind. Now, with the emergence of large scale social media websites incorporating a threaded dialogue structure, content feedback, and self-annotation (such as stance labeling), there are valuable new corpora available to researchers. In previous work, we released the INTERNET ARGUMENT CORPUS, one of the first larger scale resources available for opinion sharing dialogue. We now release the INTERNET ARGUMENT CORPUS 2.0 (IAC 2.0) in the hope that others will find it as useful as we have. The IAC 2.0 provides more data than IAC 1.0 and organizes it using an extensible, repurposable SQL schema. The database structure in conjunction with the associated code facilitates querying from and combining multiple dialogically structured data sources. The IAC 2.0 schema provides support for forum posts, quotations, markup (bold, italic, etc), and various annotations, including Stanford CoreNLP annotations. We demonstrate the generalizablity of the schema by providing code to import the ConVote corpus.
Year
Venue
Keywords
2016
LREC 2016 - TENTH INTERNATIONAL CONFERENCE ON LANGUAGE RESOURCES AND EVALUATION
dialogue,argument mining,sentiment,stance,data integration,online forums,debate
Field
DocType
Citations 
SQL,Dialogic,Social media,Computer science,Natural language processing,Artificial intelligence,Schema (psychology),Linguistics,The Internet
Conference
13
PageRank 
References 
Authors
0.59
18
4
Name
Order
Citations
PageRank
Rob Abbott12069.62
Brian Ecker2191.36
Pranav Anand326019.70
Marilyn A Walker43893418.91