Armand Joulin
15 ورقة في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Cooperative Learning of Disjoint Syntax and Semantics
2019 · arXiv (Cornell University)
Serhii Havrylov, Germán Kruszewski, Armand Joulin. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 2019.
-
Updating Pre-trained Word Vectors and Text Classifiers using Monolingual Alignment
2019 · arXiv (Cornell University)
In this paper, we focus on the problem of adapting word vector-based models to new textual data. Given a model pre-trained on large reference data, how can we adapt it to a smaller piece of …
-
Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks
2015 · arXiv (Cornell University)
One long-term goal of machine learning research is to produce methods that are applicable to reasoning and natural language, in particular building an intelligent dialogue agent. To measure progress towards that goal, we argue for …
-
Bag of Tricks for Efficient Text Classification
2017
Armand Joulin, Edouard Grave, Piotr Bojanowski, Tomas Mikolov. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers. 2017.
-
Enriching Word Vectors with Subword Information
2017 · Transactions of the Association for Computational Linguistics
Continuous word representations, trained on large unlabeled corpora are useful for many natural language processing tasks. Popular models that learn such representations ignore the morphology of words, by assigning a distinct vector to each word. …
-
Improving Neural Language Models with a Continuous Cache
2016 · arXiv (Cornell University)
We propose an extension to neural network language models to adapt their prediction to the recent history. Our model is a simplified version of memory augmented networks, which stores past hidden activations as memory and …
-
Loss in Translation: Learning Bilingual Word Mapping with a Retrieval Criterion
2018
Continuous word representations learned separately on distinct languages can be aligned so that their words become comparable in a common space. Existing works typically solve a quadratic problem to learn a orthogonal matrix aligning a …
-
Adaptive Attention Span in Transformers
2019
We propose a novel self-attention mechanism that can learn its optimal attention span. This allows us to extend significantly the maximum context size used in Transformer, while maintaining control over their memory footprint and computational …
-
Enriching Word Vectors with Subword Information
2016 · arXiv (Cornell University)
Continuous word representations, trained on large unlabeled corpora are useful for many natural language processing tasks. Popular models that learn such representations ignore the morphology of words, by assigning a distinct vector to each word. …
-
Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks
2016 · International Conference on Learning Representations
Abstract: One long-term goal of machine learning research is to produce methods that are applicable to reasoning and natural language, in particular building an intelligent dialogue agent. To measure progress towards that goal, we argue …
-
Efficient softmax approximation for GPUs
2017 · International Conference on Machine Learning
We propose an approximate strategy to efficiently train neural network based language models over very large vocabularies. Our approach, called adaptive softmax, circumvents the linear dependency on the vocabulary size by exploiting the unbalanced word …
-
Reducing Transformer Depth on Demand with Structured Dropout
2019 · arXiv (Cornell University)
Overparameterized transformer networks have obtained state of the art results in various natural language processing tasks, such as machine translation, language modeling, and question answering. These models contain hundreds of millions of parameters, necessitating a …
-
Reducing Transformer Depth on Demand with Structured Dropout
2020 · arXiv (Cornell University)
Overparametrized transformer networks have obtained state of the art results in various natural language processing tasks, such as machine translation, language modeling, and question answering. These models contain hundreds of millions of parameters, necessitating a …
-
Beyond English-Centric Multilingual Machine Translation
2020 · arXiv (Cornell University)
Existing work in translation demonstrated the potential of massively multilingual machine translation by training a single model able to translate between any pair of languages. However, much of this work is English-Centric by training only …
-
LLaMA: Open and Efficient Foundation Language Models
2023 · arXiv (Cornell University)
We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters. We train our models on trillions of tokens, and show that it is possible to train state-of-the-art models using publicly …