Trevor Cai
4 papers in the PaperMetrix corpus
Papers by this author
-
Training Compute-Optimal Large Language Models
2022 · arXiv (Cornell University)
We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the …
-
Improving language models by retrieving from trillions of tokens
2021 · arXiv (Cornell University)
We enhance auto-regressive language models by conditioning on document chunks retrieved from a large corpus, based on local similarity with preceding tokens. With a $2$ trillion token database, our Retrieval-Enhanced Transformer (RETRO) obtains comparable performance …
-
Scaling Language Models: Methods, Analysis & Insights from Training Gopher
2021 · arXiv (Cornell University)
Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world. In this paper, we present an analysis of Transformer-based language model …
-
Red Teaming Language Models with Language Models
2022
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, Geoffrey Irving. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.