Title
PDTSC 2.0 - Spoken Corpus with Rich Multi-layer Structural Annotation.
Abstract
We present a richly annotated spoken language resource, the Prague Dependency Treebank of Spoken Czech 2.0, the primary purpose of which is to serve for speech-related NLP tasks. The treebank features several novel annotation schemas close to the audio and transcript, and the morphological, syntactic and semantic annotation corresponds to the family of Prague Dependency Treebanks; it could thus be used also for linguistic studies, including comparative studies regarding text and speech. The most unique and novel feature is our approach to syntactic annotation, which differs from other similar corpora such as Treebank-3 [8] in that it does not attempt to impose syntactic structure over input, but it includes one more layer which edits the literal transcript to fluent Czech while keeping the original transcript explicitly aligned with the edited version. This allows the morphological, syntactic and semantic annotation to be deterministically and fully mapped back to the transcript and audio. It brings new possibilities for modeling morphology, syntax and semantics in spoken language – either at the original transcript with mapped annotation, or at the new layer after (automatic) editing. The corpus is publicly and freely available.
Year
Venue
Field
2017
TSD
Czech,Coreference,Annotation,Computer science,Treebank,Natural language processing,Artificial intelligence,Schema (psychology),Syntax,Spoken language,Semantics
DocType
Citations 
PageRank 
Conference
0
0.34
References 
Authors
7
6
Name
Order
Citations
PageRank
Marie Mikulová1203.97
Jiří Mírovský2223.94
Anja Nedoluzhko300.34
Petr Pajas414815.42
Jan Stepánek500.68
Jan Hajic61184109.62