conference-paper
Open access
Findings of the WMT 2018 Shared Task on Parallel Corpus Filtering
Research footprint
At a glance
- Citations
- 105
- References
- 54
- Comments
- 0
Paper overview
Abstract
We posed the shared task of assigning sentence-level quality scores for a very noisy corpus of sentence pairs crawled from the web, with the goal of sub-selecting 1% and 10% of high-quality data to be used to train machine translation systems. Seventeen participants from companies, national research labs, and universities participated in this task.
Record transparency
Publication details
- DOI
- 10.18653/v1/w18-6453
- OpenAlex
- W2902918014
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.