conference-paper وصول مفتوح

Multilingual Extension of PDTB-Style Annotation: The Case of TED Multilingual Discourse Bank

Research footprint

At a glance

الاستشهادات
19
المراجع
26
Comments
0
Paper overview

Abstract

We introduce TED-Multilingual Discourse Bank, a corpus of TED talks transcripts in 6 languages (English, German, Polish, European Portuguese, Russian and Turkish), where the ultimate aim is to provide a clearly described level of discourse structure and semantics in multiple languages.The corpus is manually annotated following the goals and principles of PDTB, involving explicit and implicit discourse connectives, entity relations, alternative lexicalizations and no relations.In the corpus, we also aim to capture the characteristics of spoken language that exist in the transcripts and adapt the PDTB scheme according to our aims; for example, we introduce hypophora.We spot other aspects of spoken discourse such as the discourse marker use of connectives to keep them distinct from their discourse connective use.TED-MDB is, to the best of our knowledge, one of the few multilingual discourse treebanks and is hoped to be a source of parallel data for contrastive linguistic analysis as well as language technology applications.We describe the corpus, the annotation procedure and provide preliminary corpus statistics.

Record transparency

Publication details

DOI
10.63317/57ndug7eq9xq
OpenAlex
W2806321739
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.