conference-paper

Usage of Morphotranslation in Data Augmentation of the Training Dataset for “TurkLang-7” Project on MT Systems Creation

  • 2022 7th International Conference on Computer Science and Engineering (UBMK)
Research footprint

At a glance

الاستشهادات
5
المراجع
0
Comments
0
Paper overview

Abstract

The paper discusses an approach to solving the problem of low-resource languages in the development of neural network machine translation systems for Turkic language pairs by artificially increasing training data in form of parallel corpora. The basis of this approach lies in usage of toolset of the multilingual portal “Turkic Morpheme” for morphological analysis and generation of word forms. The issues of creating machine translation systems for Turkic languages are considered, a hypothesis about possibility of using the presented approach to improve the results of machine learning is formulated, and the rationale for using morphological analysis methods is given. The developed algorithm for artificial augmentation of training data is presented.

Record transparency

Publication details

DOI
10.1109/ubmk55850.2022.9919552
OpenAlex
W4308095347
Document type
conference-paper
Language
EN
Source
2022 7th International Conference on Computer Science and Engineering (UBMK)
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.