Tom Hennigan
4 papers in the PaperMetrix corpus
Papers by this author
-
Training Compute-Optimal Large Language Models
2022 · arXiv (Cornell University)
We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the …
-
Improving language models by retrieving from trillions of tokens
2021 · arXiv (Cornell University)
We enhance auto-regressive language models by conditioning on document chunks retrieved from a large corpus, based on local similarity with preceding tokens. With a $2$ trillion token database, our Retrieval-Enhanced Transformer (RETRO) obtains comparable performance …
-
Scaling Language Models: Methods, Analysis & Insights from Training Gopher
2021 · arXiv (Cornell University)
Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world. In this paper, we present an analysis of Transformer-based language model …
-
Gemini: A Family of Highly Capable Multimodal Models
2023 · arXiv (Cornell University)
This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family consists of Ultra, Pro, and Nano sizes, suitable for applications ranging …