Simon Osindero
3 papers in the PaperMetrix corpus
Papers by this author
-
Training Compute-Optimal Large Language Models
2022 · arXiv (Cornell University)
We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the …
-
Improving language models by retrieving from trillions of tokens
2021 · arXiv (Cornell University)
We enhance auto-regressive language models by conditioning on document chunks retrieved from a large corpus, based on local similarity with preceding tokens. With a $2$ trillion token database, our Retrieval-Enhanced Transformer (RETRO) obtains comparable performance …
-
Scaling Language Models: Methods, Analysis & Insights from Training Gopher
2021 · arXiv (Cornell University)
Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world. In this paper, we present an analysis of Transformer-based language model …