Erich Elsen
5 papers in the PaperMetrix corpus
Papers by this author
-
DSD: Regularizing Deep Neural Networks with Dense-Sparse-Dense Training Flow.
2016 · arXiv (Cornell University)
Modern deep neural networks have a large number of parameters, making them very powerful machine learning systems. A critical issue for training such large networks on large-scale data-sets is to prevent overfitting while at the …
-
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
2015 · arXiv (Cornell University)
We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech--two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning …
-
Training Compute-Optimal Large Language Models
2022 · arXiv (Cornell University)
We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the …
-
Improving language models by retrieving from trillions of tokens
2021 · arXiv (Cornell University)
We enhance auto-regressive language models by conditioning on document chunks retrieved from a large corpus, based on local similarity with preceding tokens. With a $2$ trillion token database, our Retrieval-Enhanced Transformer (RETRO) obtains comparable performance …
-
Scaling Language Models: Methods, Analysis & Insights from Training Gopher
2021 · arXiv (Cornell University)
Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world. In this paper, we present an analysis of Transformer-based language model …