Holger Schwenk
13 papers in the PaperMetrix corpus
Papers by this author
-
Empirical Use of Information Retrieval to Build Synthetic Data for SMT Domain Adaptation
2016 · IEEE/ACM Transactions on Audio Speech and Language Processing
In this paper, we present information retrieval as a powerful tool for addressing an imperative problem in the field of statistical machine translation, i.e., improving translation quality when not enough parallel corpora are available. We …
-
Very Deep Convolutional Networks for Text Classification
2017
The dominant approach for many NLP tasks are recurrent neural networks, in particular LSTMs, and convolutional neural networks. However, these architectures are rather shallow in comparison to the deep convolutional networks which have pushed the …
-
Learning Joint Multilingual Sentence Representations with Neural Machine Translation
2017
In this paper, we use the framework of neural machine translation to learn joint sentence representations across six very different languages. Our aim is that a representation which is independent of the language, is likely …
-
MLQA: Evaluating Cross-lingual Extractive Question Answering
2020
Question answering (QA) models have shown rapid progress enabled by the availability of large, high-quality benchmark datasets. Such annotated datasets are difficult and costly to collect, and rarely exist in languages other than English, making …
-
On Using Monolingual Corpora in Neural Machine Translation
2015 · HAL (Le Centre pour la Communication Scientifique Directe)
Recent work on end-to-end neural network-based architectures for machine translation has shown promising results for En-Fr and En-De translation. Arguably, one of the major factors behind this success has been the availability of high quality …
-
Very Deep Convolutional Networks for Natural Language Processing.
2016 · arXiv (Cornell University)
The dominant approach for many NLP tasks are recurrent neural networks, in particular LSTMs, and convolutional neural networks. However, these architectures are rather shallow in comparison to the deep convolutional networks which are very successful …
-
XNLI: Evaluating Cross-lingual Sentence Representations
2018 · arXiv (Cornell University)
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, Veselin Stoyanov. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018.
-
Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and Beyond
2019 · Transactions of the Association for Computational Linguistics
We introduce an architecture to learn joint multilingual sentence representations for 93 languages, belonging to more than 30 different families and written in 28 different scripts. Our system uses a single BiLSTM encoder with a …
-
WikiMatrix: Mining 135M Parallel Sentences in 1620 Language Pairs from Wikipedia
2021
Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, Francisco Guzmán. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
-
Margin-based Parallel Corpus Mining with Multilingual Sentence Embeddings
2018 · arXiv (Cornell University)
Machine translation is highly sensitive to the size and quality of the training data, which has led to an increasing interest in collecting and filtering large parallel corpora. In this paper, we propose a new …
-
Supervised Learning of Universal Sentence Representations from Natural\n Language Inference Data
2017 · arXiv (Cornell University)
Many modern NLP systems rely on word embeddings, previously trained in an\nunsupervised manner on large corpora, as base features. Efforts to obtain\nembeddings for larger chunks of text, such as sentences, have however not been\nso successful. …
-
Beyond English-Centric Multilingual Machine Translation
2020 · arXiv (Cornell University)
Existing work in translation demonstrated the potential of massively multilingual machine translation by training a single model able to translate between any pair of languages. However, much of this work is English-Centric by training only …
-
No Language Left Behind: Scaling Human-Centered Machine Translation
2022 · arXiv (Cornell University)
Driven by the goal of eradicating language barriers on a global scale, machine translation has solidified itself as a key focus of artificial intelligence research today. However, such efforts have coalesced around a small subset …