PDTSC 2.0 - Spoken Corpus with Rich Multi-layer Structural Annotation. - Citegraph

Paper Info

Title
PDTSC 2.0 - Spoken Corpus with Rich Multi-layer Structural Annotation.

Abstract
We present a richly annotated spoken language resource, the Prague Dependency Treebank of Spoken Czech 2.0, the primary purpose of which is to serve for speech-related NLP tasks. The treebank features several novel annotation schemas close to the audio and transcript, and the morphological, syntactic and semantic annotation corresponds to the family of Prague Dependency Treebanks; it could thus be used also for linguistic studies, including comparative studies regarding text and speech. The most unique and novel feature is our approach to syntactic annotation, which differs from other similar corpora such as Treebank-3 [8] in that it does not attempt to impose syntactic structure over input, but it includes one more layer which edits the literal transcript to fluent Czech while keeping the original transcript explicitly aligned with the edited version. This allows the morphological, syntactic and semantic annotation to be deterministically and fully mapped back to the transcript and audio. It brings new possibilities for modeling morphology, syntax and semantics in spoken language – either at the original transcript with mapped annotation, or at the new layer after (automatic) editing. The corpus is publicly and freely available.

Year	Venue	Field
2017	TSD	Czech,Coreference,Annotation,Computer science,Treebank,Natural language processing,Artificial intelligence,Schema (psychology),Syntax,Spoken language,Semantics
DocType	Citations	PageRank
Conference	0	0.34
References	Authors
7	6

Authors (6 rows)

Cited by (0 rows)

References (7 rows)

Name	Order	Citations	PageRank
Marie Mikulová	1	20	3.97
Jiří Mírovský	2	22	3.94
Anja Nedoluzhko	3	0	0.34
Petr Pajas	4	148	15.42
Jan Stepánek	5	0	0.68
Jan Hajic	6	1184	109.62

1