conference-paper Open access

A Multilingual Multiway Evaluation Data Set for Structured Document Translation of Asian Languages

Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Abstract

Translation of structured content is an important application of machine translation, but the scarcity of evaluation data sets, especially for Asian languages, limits progress.In this paper we present a novel multilingual multiway evaluation data set for the translation of structured documents of the Asian languages Japanese, Korean and Chinese.We describe the data set, its creation process and important characteristics, followed by establishing and evaluating baselines using the direct translation as well as detag-project approaches.Our data set is well suited for multilingual evaluation, and it contains richer annotation tag sets than existing data sets.Our results show that massively multilingual translation models like M2M-100 and mBART-50 perform surprisingly well despite not being explicitly trained to handle structured content.The data set described in this paper and used in our experiments is released publicly.

Record transparency

Publication details

DOI
10.18653/v1/2022.findings-aacl.23
OpenAlex
W4404783710
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.