Mikel L. Forcada
4 papers in the PaperMetrix corpus
Papers by this author
-
Stand-off Annotation of Web Content as a Legally Safer Alternative to Crawling for Distribution
2016 · RUA, Repositorio Institucional de la Universidad de Alicante (Universidad de Alicante)
Sentence-aligned web-crawled parallel text or bitext is frequently used to train statistical machine translation systems. To that end, web-crawled sentence-aligned bitext sets are sometimes made publicly available and distributed by translation technologies practitioners. Contrary to …
-
A multi-source approach for Breton–French hybrid machine translation
2020 · Zenodo (CERN European Organization for Nuclear Research)
Corpus-based approaches to machine translation (MT) have difficulties when the amount of parallel corpora to use for training is scarce, especially if the languages involved in the translation are highly inflected. This problem can be …
-
Findings of the WMT 2018 Shared Task on Parallel Corpus Filtering
2018
We posed the shared task of assigning sentence-level quality scores for a very noisy corpus of sentence pairs crawled from the web, with the goal of sub-selecting 1% and 10% of high-quality data to be …
-
ParaCrawl: Web-Scale Acquisition of Parallel Corpora
2020
Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strelec, Brian Thompson, …