conference-paper

Usage of Morphotranslation in Data Augmentation of the Training Dataset for “TurkLang-7” Project on MT Systems Creation

  • 2022 7th International Conference on Computer Science and Engineering (UBMK)
Research footprint

At a glance

Citations
5
References
0
Comments
0
Paper overview

Öz

The paper discusses an approach to solving the problem of low-resource languages in the development of neural network machine translation systems for Turkic language pairs by artificially increasing training data in form of parallel corpora. The basis of this approach lies in usage of toolset of the multilingual portal “Turkic Morpheme” for morphological analysis and generation of word forms. The issues of creating machine translation systems for Turkic languages are considered, a hypothesis about possibility of using the presented approach to improve the results of machine learning is formulated, and the rationale for using morphological analysis methods is given. The developed algorithm for artificial augmentation of training data is presented.

Record transparency

Publication details

DOI
10.1109/ubmk55850.2022.9919552
OpenAlex
W4308095347
Document type
conference-paper
Language
EN
Source
2022 7th International Conference on Computer Science and Engineering (UBMK)
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.