preprint Open access

Evaluating Contextualized Language Models for Hungarian

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
3
References
4
Comments
0
Paper overview

Abstract

We present an extended comparison of contextualized language models for Hungarian. We compare huBERT, a Hungarian model against 4 multilingual models including the multilingual BERT model. We evaluate these models through three tasks, morphological probing, POS tagging and NER. We find that huBERT works better than the other models, often by a large margin, particularly near the global optimum (typically at the middle layers). We also find that huBERT tends to generate fewer subwords for one word and that using the last subword for token-level tasks is generally a better choice than using the first one.

Record transparency

Publication details

DOI
10.48550/arxiv.2102.10848
OpenAlex
W3131580605
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.