Researcher profile

Oriol Vinyals

19 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Emergent Abilities of Large Language Models

    2022 · arXiv (Cornell University)

    Scaling up language models has been shown to predictably improve performance and sample efficiency on a wide range of downstream tasks. This paper instead discusses an unpredictable phenomenon that we refer to as emergent abilities …

  2. Fantastic Gains and Where to Find Them: On the Existence and Prospect of General Knowledge Transfer between Any Pretrained Model

    2023 · arXiv (Cornell University)

    Training deep networks requires various design decisions regarding for instance their architecture, data augmentation, or optimization. In this work, we find these training variations to result in networks learning unique feature sets from the data. …

  3. Pointer Networks

    2015 · arXiv (Cornell University)

    We introduce a new neural architecture to learn the conditional probability of an output sequence with elements that are discrete tokens corresponding to positions in an input sequence. Such problems cannot be trivially addressed by …

  4. A Neural Conversational Model

    2015 · arXiv (Cornell University)

    Conversational modeling is an important task in natural language understanding and machine intelligence. Although previous approaches exist, they are often restricted to specific domains (e.g., booking an airline ticket) and require hand-crafted rules. In this …

  5. Show and tell: A neural image caption generator

    2015

    Automatically describing the content of an image is a fundamental problem in artificial intelligence that connects computer vision and natural language processing. In this paper, we present a generative model based on a deep recurrent …

  6. Addressing the Rare Word Problem in Neural Machine Translation

    2015

    Thang Luong, Ilya Sutskever, Quoc Le, Oriol Vinyals, Wojciech Zaremba. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long …

  7. Multi-task Sequence to Sequence Learning

    2015 · arXiv (Cornell University)

    Sequence to sequence learning has recently emerged as a new paradigm in supervised learning. To date, most of its applications focused on only one task and not much work explored this framework for multiple tasks. …

  8. Order Matters: Sequence to sequence for sets

    2015 · arXiv (Cornell University)

    Sequences have become first class citizens in supervised learning thanks to the resurgence of recurrent neural networks. Many complex tasks that require mapping from or to a sequence of observations can now be formulated with …

  9. Generating Sentences from a Continuous Space

    2016

    The standard recurrent neural network language model (rnnlm) generates sentences one word at a time and does not work from an explicit global sentence representation. In this work, we introduce and study an rnn-based variational …

  10. Sentence Compression by Deletion with LSTMs

    2015

    We present an LSTM approach to deletion-based sentence compression where the task is to translate a sentence into a sequence of zeros and ones, corresponding to token deletion decisions. We demonstrate that even the most …

  11. Exploring the Limits of Language Modeling

    2016 · arXiv (Cornell University)

    In this work we explore recent advances in Recurrent Neural Networks for large scale Language Modeling, a task central to language understanding. We extend current models to deal with two key challenges present in this …

  12. Listen, attend and spell: A neural network for large vocabulary conversational speech recognition

    2016

    We present Listen, Attend and Spell (LAS), a neural speech recognizer that transcribes speech utterances directly to characters without pronunciation models, HMMs or other components of traditional speech recognizers. In LAS, the neural network architecture …

  13. Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation

    2016 · arXiv (Cornell University)

    Neural Machine Translation (NMT) is an end-to-end learning approach for automated translation, with the potential to overcome many of the weaknesses of conventional phrase-based translation systems. Unfortunately, NMT systems are known to be computationally expensive …

  14. Universal Transformers

    2018 · arXiv (Cornell University)

    Recurrent neural networks (RNNs) sequentially process data by updating their state with each new data point, and have long been the de facto choice for sequence modeling tasks. However, their inherently sequential computation makes them …

  15. Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks

    2015 · arXiv (Cornell University)

    Recurrent Neural Networks can be trained to produce sequences of tokens given some input, as exemplified by recent results in machine translation and image captioning. The current approach to training them consists of maximizing the …

  16. Multilingual Language Processing From Bytes

    2016

    Dan Gillick, Cliff Brunk, Oriol Vinyals, Amarnag Subramanya. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.

  17. Training Compute-Optimal Large Language Models

    2022 · arXiv (Cornell University)

    We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the …

  18. Improving language models by retrieving from trillions of tokens

    2021 · arXiv (Cornell University)

    We enhance auto-regressive language models by conditioning on document chunks retrieved from a large corpus, based on local similarity with preceding tokens. With a $2$ trillion token database, our Retrieval-Enhanced Transformer (RETRO) obtains comparable performance …

  19. Scaling Language Models: Methods, Analysis & Insights from Training Gopher

    2021 · arXiv (Cornell University)

    Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world. In this paper, we present an analysis of Transformer-based language model …