IIT Delhi Dialogue Corpus: A Quantitative Analysis of a Spoken Corpus of Hindi
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Öz
We present our effort to create a dialogue corpus for Hindi with the aim of under-standing (a) the nature of linguistic utterances during naturalistic dialogue, (b)what these linguistic patterns tell us about the cognitive processes/constraints that affect production and comprehension during dialogue, and (c) how do such processes/constraints differ from written text. We discuss the procedure and pipeline employed to create two sets of spoken data -- telephonic conversation data, and face-to-face (task-oriented) conversation data. At the lexical level, the data has been annotated for information such as disfluencies, code-switching, etc., and at the syntactic level for part-of-speech tags and dependency relations.We present a preliminary analysis of the created dialogue data and compare it with a written text to discuss the usefulness and implications of this resource for psycholinguistic research.
Publication details
- DOI
- 10.31219/osf.io/tc7f5_v1
- OpenAlex
- W4407896186
- Document type
- preprint
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.