Édouard Grave
15 papers in the PaperMetrix corpus
Papers by this author
-
Misspelling Oblivious Word Embeddings
2019
Aleksandra Piktus, Necati Bora Edizel, Piotr Bojanowski, Edouard Grave, Rui Ferreira, Fabrizio Silvestri. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long …
-
Updating Pre-trained Word Vectors and Text Classifiers using Monolingual Alignment
2019 · arXiv (Cornell University)
In this paper, we focus on the problem of adapting word vector-based models to new textual data. Given a model pre-trained on large reference data, how can we adapt it to a smaller piece of …
-
Bag of Tricks for Efficient Text Classification
2017
Armand Joulin, Edouard Grave, Piotr Bojanowski, Tomas Mikolov. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers. 2017.
-
Enriching Word Vectors with Subword Information
2017 · Transactions of the Association for Computational Linguistics
Continuous word representations, trained on large unlabeled corpora are useful for many natural language processing tasks. Popular models that learn such representations ignore the morphology of words, by assigning a distinct vector to each word. …
-
Improving Neural Language Models with a Continuous Cache
2016 · arXiv (Cornell University)
We propose an extension to neural network language models to adapt their prediction to the recent history. Our model is a simplified version of memory augmented networks, which stores past hidden activations as memory and …
-
Loss in Translation: Learning Bilingual Word Mapping with a Retrieval Criterion
2018
Continuous word representations learned separately on distinct languages can be aligned so that their words become comparable in a common space. Existing works typically solve a quadratic problem to learn a orthogonal matrix aligning a …
-
Adaptive Attention Span in Transformers
2019
We propose a novel self-attention mechanism that can learn its optimal attention span. This allows us to extend significantly the maximum context size used in Transformer, while maintaining control over their memory footprint and computational …
-
Enriching Word Vectors with Subword Information
2016 · arXiv (Cornell University)
Continuous word representations, trained on large unlabeled corpora are useful for many natural language processing tasks. Popular models that learn such representations ignore the morphology of words, by assigning a distinct vector to each word. …
-
Efficient softmax approximation for GPUs
2017 · International Conference on Machine Learning
We propose an approximate strategy to efficiently train neural network based language models over very large vocabularies. Our approach, called adaptive softmax, circumvents the linear dependency on the vocabulary size by exploiting the unbalanced word …
-
Can You Tell Me How to Get Past Sesame Street? Sentence-Level Pretraining Beyond Language Modeling
2019
Alex Wang, Jan Hula, Patrick Xia, Raghavendra Pappagari, R. Thomas McCoy, Roma Patel, Najoung Kim, Ian Tenney, Yinghui Huang, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, Samuel R. Bowman. …
-
Reducing Transformer Depth on Demand with Structured Dropout
2019 · arXiv (Cornell University)
Overparameterized transformer networks have obtained state of the art results in various natural language processing tasks, such as machine translation, language modeling, and question answering. These models contain hundreds of millions of parameters, necessitating a …
-
Unsupervised Cross-lingual Representation Learning at Scale
2020
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, Veselin Stoyanov. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
-
Reducing Transformer Depth on Demand with Structured Dropout
2020 · arXiv (Cornell University)
Overparametrized transformer networks have obtained state of the art results in various natural language processing tasks, such as machine translation, language modeling, and question answering. These models contain hundreds of millions of parameters, necessitating a …
-
Beyond English-Centric Multilingual Machine Translation
2020 · arXiv (Cornell University)
Existing work in translation demonstrated the potential of massively multilingual machine translation by training a single model able to translate between any pair of languages. However, much of this work is English-Centric by training only …
-
LLaMA: Open and Efficient Foundation Language Models
2023 · arXiv (Cornell University)
We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters. We train our models on trillions of tokens, and show that it is possible to train state-of-the-art models using publicly …