article
Open access
Automatic Diacritization as Prerequisite Towards the Automatic Generation of Arabic Lexical Recognition Tests
Research footprint
At a glance
- Citations
- 1
- References
- 24
- Comments
- 0
Paper overview
Abstract
The automatic generation of Arabic lexical recognition tests entails several NLP challenges, including corpus linguistics, automatic diacritization, lemmatization and language modeling. Here, we only address the problem of automatic diacritization, a step that paves the road for the automatic generation of Arabic LRTs. We conduct a comparative study between the available tools for diacritization (Farasa and Madamira) and a strong baseline. We evaluate the error rates for these systems using a set of publicly available (almost) fully diacritized corpora, but in a relaxed evaluation mode to ensure fair comparison. Farasa outperforms Madamira and the baseline under all conditions.
Record transparency
Publication details
- DOI
- 10.17185/duepublico/72018
- OpenAlex
- W2991195880
- Document type
- article
- Language
- EN
- Source
- DuEPublico (University of Duisburg-Essen)
- Last metadata update
Comments
Log in to join the discussion.