conference-paper وصول مفتوح

Document-aligned Japanese-English Conversation Parallel Corpus

Research footprint

At a glance

الاستشهادات
0
المراجع
18
Comments
0
Paper overview

Abstract

Sentence-level (SL) machine translation (MT) has reached acceptable quality for many highresourced languages, but not document-level (DL) MT, which is difficult to 1) train with little amount of DL data; and 2) evaluate, as the main methods and data sets focus on SL evaluation.To address the first issue, we present a document-aligned Japanese-English conversation corpus, including balanced, high-quality business conversation data for tuning and testing.As for the second issue, we manually identify the main areas where SL MT fails to produce adequate translations in lack of context.We then create an evaluation set where these phenomena are annotated to alleviate automatic evaluation of DL systems.We train MT models using our corpus to demonstrate how using context leads to improvements.

Record transparency

Publication details

DOI
10.18653/v1/2020.wmt-1.74
OpenAlex
W3121059898
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.