article

Information Dropping Data Augmentation for Machine Translation Quality Estimation

  • IEEE/ACM Transactions on Audio Speech and Language Processing
  • Institute of Electrical and Electronics Engineers
Research footprint

At a glance

Citations
5
References
62
Comments
0
Paper overview

Abstract

Machine translation quality estimation (QE) refers to the quality assessment of machine translations without a given reference translation. Supervised QE models based on neural networks have achieved state-of-the-art results. But this method requires large-scale training data, which requires bilingual experts to create high-quality labels. This is often very costly. Therefore, we propose a sentence-level machine translation QE data augmentation method based on information dropping. Firstly, we calculate the subwords information of the target translation based on the conditional language model. Subsequently, some subwords in the target translation are randomly deleted or replaced. We obtain the pseudo quality score by calculating the remaining information. Finally, the original and augmented data are combined to train the final model. This pseudo-data generation method based on information dropping strategy enables us to obtain more faithful and diverse training samples without requiring additional corpus resources. Experimental results show that we improve the correlation with human judgment by an average of 5.96% in the seven translation directions of the MLQE-PE dataset, while improving the model's robustness to low adequacy samples. In addition, the method does not require any modifications to the model architecture.

Record transparency

Publication details

DOI
10.1109/taslp.2024.3380996
OpenAlex
W4393078878
Document type
article
Language
EN
Source
IEEE/ACM Transactions on Audio Speech and Language Processing
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.