Researcher profile

Édouard Grave

15 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Misspelling Oblivious Word Embeddings

    2019

    Aleksandra Piktus, Necati Bora Edizel, Piotr Bojanowski, Edouard Grave, Rui Ferreira, Fabrizio Silvestri. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long …

  2. Updating Pre-trained Word Vectors and Text Classifiers using Monolingual Alignment

    2019 · arXiv (Cornell University)

    In this paper, we focus on the problem of adapting word vector-based models to new textual data. Given a model pre-trained on large reference data, how can we adapt it to a smaller piece of …

  3. Bag of Tricks for Efficient Text Classification

    2017

    Armand Joulin, Edouard Grave, Piotr Bojanowski, Tomas Mikolov. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers. 2017.

  4. Enriching Word Vectors with Subword Information

    2017 · Transactions of the Association for Computational Linguistics

    Continuous word representations, trained on large unlabeled corpora are useful for many natural language processing tasks. Popular models that learn such representations ignore the morphology of words, by assigning a distinct vector to each word. …

  5. Improving Neural Language Models with a Continuous Cache

    2016 · arXiv (Cornell University)

    We propose an extension to neural network language models to adapt their prediction to the recent history. Our model is a simplified version of memory augmented networks, which stores past hidden activations as memory and …

  6. Loss in Translation: Learning Bilingual Word Mapping with a Retrieval Criterion

    2018

    Continuous word representations learned separately on distinct languages can be aligned so that their words become comparable in a common space. Existing works typically solve a quadratic problem to learn a orthogonal matrix aligning a …

  7. Adaptive Attention Span in Transformers

    2019

    We propose a novel self-attention mechanism that can learn its optimal attention span. This allows us to extend significantly the maximum context size used in Transformer, while maintaining control over their memory footprint and computational …

  8. Enriching Word Vectors with Subword Information

    2016 · arXiv (Cornell University)

    Continuous word representations, trained on large unlabeled corpora are useful for many natural language processing tasks. Popular models that learn such representations ignore the morphology of words, by assigning a distinct vector to each word. …

  9. Efficient softmax approximation for GPUs

    2017 · International Conference on Machine Learning

    We propose an approximate strategy to efficiently train neural network based language models over very large vocabularies. Our approach, called adaptive softmax, circumvents the linear dependency on the vocabulary size by exploiting the unbalanced word …

  10. Can You Tell Me How to Get Past Sesame Street? Sentence-Level Pretraining Beyond Language Modeling

    2019

    Alex Wang, Jan Hula, Patrick Xia, Raghavendra Pappagari, R. Thomas McCoy, Roma Patel, Najoung Kim, Ian Tenney, Yinghui Huang, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, Samuel R. Bowman. …

  11. Reducing Transformer Depth on Demand with Structured Dropout

    2019 · arXiv (Cornell University)

    Overparameterized transformer networks have obtained state of the art results in various natural language processing tasks, such as machine translation, language modeling, and question answering. These models contain hundreds of millions of parameters, necessitating a …

  12. Unsupervised Cross-lingual Representation Learning at Scale

    2020

    Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, Veselin Stoyanov. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.

  13. Reducing Transformer Depth on Demand with Structured Dropout

    2020 · arXiv (Cornell University)

    Overparametrized transformer networks have obtained state of the art results in various natural language processing tasks, such as machine translation, language modeling, and question answering. These models contain hundreds of millions of parameters, necessitating a …

  14. Beyond English-Centric Multilingual Machine Translation

    2020 · arXiv (Cornell University)

    Existing work in translation demonstrated the potential of massively multilingual machine translation by training a single model able to translate between any pair of languages. However, much of this work is English-Centric by training only …

  15. LLaMA: Open and Efficient Foundation Language Models

    2023 · arXiv (Cornell University)

    We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters. We train our models on trillions of tokens, and show that it is possible to train state-of-the-art models using publicly …