article Open access

Low-Resource Noisy Transliteration Normalization Using Large-Scale Language Model

  • IEEE Access
  • Institute of Electrical and Electronics Engineers
Research footprint

At a glance

Citations
1
References
23
Comments
0
Paper overview

Abstract

Transliteration normalization is a crucial task for low-resource languages, particularly for Mongolian, where noisy text from social media presents significant challenges. The frequent use of non-standard transliteration can contribute to the gradual erosion of linguistic knowledge, particularly among young users, making it harder to maintain proficiency in their native language. Therefore, developing robust methods for normalizing such text is essential. In this paper, we propose a novel approach leveraging large-scale neural models, specifically GPT-2, to normalize noisy transliterated Mongolian text. Our study explores a data-driven approach, including word pairs, sentence pairs, and synthetic data, to enhance model performance. To further improve accuracy, we introduce a post-processing module that integrates Edit Distance-based corrections with a context-aware ranking mechanism using the Mongolian BERT model. Experimental results demonstrate that our approach (M10: 16.42%) improves overall accuracy by approximately 4.91%, while achieving a 10.44% increase in out-of-vocabulary (OOV) word normalization compared to baseline models. Our proposed approach demonstrates effectiveness in normalizing noisy transliterated text under low-resource conditions.

Record transparency

Publication details

DOI
10.1109/access.2025.3574933
OpenAlex
W4410852748
Document type
article
Language
EN
Source
IEEE Access
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.