conference-paper Open access

Romanian Lexical Resources Interconnection

  • Procedia Computer Science
  • Elsevier BV
Research footprint

At a glance

Citations
2
References
14
Comments
0
Paper overview

Öz

Great efforts are being made to increase the utility of linguistic resources in applications related to language processing by interconnecting them. This study is focused on the Romanian language, which is a young language in this domain, for which it has to be made big steps until it is considered well resourced - both qualitatively and quantitatively. The CoRoLa corpus has been developed for 4 years and since 2017 it is visible for research. It cumulates about one billion words, annotated (tokenized, lemmatized, morphologically and syntactically processed). Other two resources used in this study are the eDTLR (the electronic version of the Thesaurus Dictionary of the Romanian Language) and Romanian WordNet. In this study it is described a technology that unifies these resources and brings them to a common standard, which leads the way for an efficient coupling of these resources. With this new obtained standardized resource, a series of test and case studies were performed, with the benefit of a much more diverse exposure on using the Romanian language.

Record transparency

Publication details

DOI
10.1016/j.procs.2021.08.075
OpenAlex
W3202620303
Document type
conference-paper
Language
EN
Source
Procedia Computer Science
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.