Enriching and Increasing the Usability of Lexicographical Data for Less-Resourced
At a glance
- Citations
- 0
- References
- 9
- Comments
- 0
Öz
This paper presents a use case for enriching lexicographical data for less-resourced languages employing the CLARIN infrastructure. Newly prepared lexicographical data sets for under-resourced Bantu languages spoken in southern regions of the African continent form the basis of the presented work. These datasets have been made digitally available using well-established standards of the Linguistic Linked Open Data (LLOD) community. To overcome the insufficient amount of freely available reference material, a crowdsourcing web portal for collecting textual data for less-resourced languages has been created and incorporated into the CLARIN infrastructure. Using this portal, the number of available text resources for the respective languages was significantly increased in a community effort. The collected content is used to enrich lexicographical data with real-world samples to increase the usability of the entire resource.
Publication details
- DOI
- 10.3384/ecp2020172004
- OpenAlex
- W3039545895
- Document type
- conference-paper
- Language
- EN
- Source
- Linköping electronic conference proceedings
- Last metadata update
Comments
Oturum Açın to join the discussion.