article Open access

Automatic Diacritization as Prerequisite Towards the Automatic Generation of Arabic Lexical Recognition Tests

  • DuEPublico (University of Duisburg-Essen)
  • University of Duisburg-Essen
Research footprint

At a glance

Citations
1
References
24
Comments
0
Paper overview

Abstract

The automatic generation of Arabic lexical recognition tests entails several NLP challenges, including corpus linguistics, automatic diacritization, lemmatization and language modeling. Here, we only address the problem of automatic diacritization, a step that paves the road for the automatic generation of Arabic LRTs. We conduct a comparative study between the available tools for diacritization (Farasa and Madamira) and a strong baseline. We evaluate the error rates for these systems using a set of publicly available (almost) fully diacritized corpora, but in a relaxed evaluation mode to ensure fair comparison. Farasa outperforms Madamira and the baseline under all conditions.

Record transparency

Publication details

DOI
10.17185/duepublico/72018
OpenAlex
W2991195880
Document type
article
Language
EN
Source
DuEPublico (University of Duisburg-Essen)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.