preprint وصول مفتوح

Correct Me If You Can: Learning from Error Corrections and Markings

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

الاستشهادات
3
المراجع
33
Comments
0
Paper overview

Abstract

Sequence-to-sequence learning involves a trade-off between signal strength and annotation cost of training data. For example, machine translation data range from costly expert-generated translations that enable supervised learning, to weak quality-judgment feedback that facilitate reinforcement learning. We present the first user study on annotation cost and machine learnability for the less popular annotation mode of error markings. We show that error markings for translations of TED talks from English to German allow precise credit assignment while requiring significantly less human effort than correcting/post-editing, and that error-marked data can be used successfully to fine-tune neural machine translation models.

Record transparency

Publication details

DOI
10.48550/arxiv.2004.11222
OpenAlex
W3018072206
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.