Mikel Artetxe
11 ورقة في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Translation Artifacts in Cross-lingual Transfer Learning
2020 · Communities in ADDI (Universidad del Pais Vasco)
Both human and machine translation play a central role in cross-lingual transfer learning: many multilingual datasets have been created through professional translation services, and using machine translation to translate either the test set or the …
-
Lifting the Curse of Multilinguality by Pre-training Modular Transformers
2022 · arXiv (Cornell University)
Multilingual pre-trained models are known to suffer from the curse of multilinguality, which causes per-language performance to drop as they cover more languages. We address this issue by introducing language-specific modules, which allows us to …
-
Learning principled bilingual mappings of word embeddings while preserving monolingual invariance
2016
Mapping word embeddings of different languages into a single space has multiple applications. In order to map from a source space into a target space, a common approach is to learn a linear mapping that …
-
Learning bilingual word embeddings with (almost) no bilingual data
2017
Most methods to learn bilingual word embeddings rely on large parallel corpora, which is difficult to obtain for most language pairs. This has motivated an active research line to relax this requirement, with methods that …
-
Generalizing and Improving Bilingual Word Embedding Mappings with a Multi-Step Framework of Linear Transformations
2018 · Proceedings of the AAAI Conference on Artificial Intelligence
Using a dictionary to map independently trained word embeddings to a shared space has shown to be an effective approach to learn bilingual word embeddings. In this work, we propose a multi-step framework of linear …
-
Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and Beyond
2019 · Transactions of the Association for Computational Linguistics
We introduce an architecture to learn joint multilingual sentence representations for 93 languages, belonging to more than 30 different families and written in 28 different scripts. Our system uses a single BiLSTM encoder with a …
-
Margin-based Parallel Corpus Mining with Multilingual Sentence Embeddings
2018 · arXiv (Cornell University)
Machine translation is highly sensitive to the size and quality of the training data, which has led to an increasing interest in collecting and filtering large parallel corpora. In this paper, we propose a new …
-
A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings
2018 · Communities in ADDI (University of the Basque Country)
Recent work has managed to learn cross-lingual word embeddings without parallel data by mapping monolingual embeddings to a shared space through adversarial training. However, their evaluation has focused on favorable conditions, using comparable corpora or …
-
On the Cross-lingual Transferability of Monolingual Representations
2020
State-of-the-art unsupervised multilingual models (e.g., multilingual BERT) have been shown to generalize in a zero-shot cross-lingual setting. This generalization ability has been attributed to the use of a shared subword vocabulary and joint training across …
-
Analysing Off-The-Shelf Options for Question Answering with Portuguese FAQs
2022 · DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)
Following the current interest in developing automatic question answering systems, we analyse alternative approaches for finding suitable answers from a list of Frequently Asked Questions (FAQs), in Portuguese. These rely on different technologies, some more …
-
Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?
2022
Large language models (LMs) are able to in-context learn—perform a new task via inference alone by conditioning on a few input-label pairs (demonstrations) and making predictions for new inputs. However, there has been little understanding …