Researcher profile

Karen Simonyan

4 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Neural Machine Translation in Linear Time

    2016 · arXiv (Cornell University)

    We present a novel neural network for processing sequences. The ByteNet is a one-dimensional convolutional neural network that is composed of two parts, one to encode the source sequence and the other to decode the …

  2. Training Compute-Optimal Large Language Models

    2022 · arXiv (Cornell University)

    We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the …

  3. Improving language models by retrieving from trillions of tokens

    2021 · arXiv (Cornell University)

    We enhance auto-regressive language models by conditioning on document chunks retrieved from a large corpus, based on local similarity with preceding tokens. With a $2$ trillion token database, our Retrieval-Enhanced Transformer (RETRO) obtains comparable performance …

  4. Scaling Language Models: Methods, Analysis & Insights from Training Gopher

    2021 · arXiv (Cornell University)

    Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world. In this paper, we present an analysis of Transformer-based language model …