Malayalam to English Named Entity Transliteration using Attention based BiLSTM
At a glance
- Citations
- 0
- References
- 10
- Comments
- 0
Abstract
This research introduces an approach to Malayalam-English named entity transliteration using a BiLSTM with attention mechanism and compares its perfromance with basic LSTM and BiLSTM frame-works. A selected subset of 27 million named entity parallel corpus is employed for training the translit-eration model. The trained model exhibits promising results, achieving an accuracy of 53% on a held-out test set, with an impressively low character error rate (CER) of 7.7%. Notably, our model demonstrates a high degree of generalization beyond named entities. In testing its capabilities on non-named entities. The character tokenizer applied in this model has shown surprising proficiency in transliterating regular words in Malayalam, including single morpheme words, as well as those with inflections and agglutinations. This research underscores the potential of utilizing spe-cialized parallel corpora for enhancing transliteration models. Furthermore, the flexibility of the proposed architecture allows for training reverse transliteration, converting English names to Malayalam script. The findings presented here contribute valuable insights to the broader domain of cross-language information processing and pave the way for future improvements in the transliteration of Indian languages.
Publication details
- DOI
- 10.1109/raics61201.2024.10690040
- OpenAlex
- W4403024062
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.