conference-paper Open access

New Inflectional Lexicons and Training Corpora for Improved Morphosyntactic Annotation of Croatian and Serbian

Research footprint

At a glance

Citations
57
References
14
Comments
0
Paper overview

Abstract

In this paper we present newly developed inflectional lexcions and manually annotated corpora of Croatian and Serbian.We introduce hrLex and srLex-two freely available inflectional lexicons of Croatian and Serbian-and describe the process of building these lexicons, supported by supervised machine learning techniques for lemma and paradigm prediction.Furthermore, we introduce hr500k, a manually annotated corpus of Croatian, 500 thousand tokens in size.We showcase the three newly developed resources on the task of morphosyntactic annotation of both languages by using a recently developed CRF tagger.We achieve best results yet reported on the task for both languages, beating the HunPos baseline trained on the same datasets by a wide margin.

Record transparency

Publication details

DOI
10.63317/3pnx6nwno2ic
OpenAlex
W2573540318
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.