Sharan Narang
8 papers in the PaperMetrix corpus
Papers by this author
-
DSD: Regularizing Deep Neural Networks with Dense-Sparse-Dense Training Flow.
2016 · arXiv (Cornell University)
Modern deep neural networks have a large number of parameters, making them very powerful machine learning systems. A critical issue for training such large networks on large-scale data-sets is to prevent overfitting while at the …
-
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
2015 · arXiv (Cornell University)
We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech--two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning …
-
Deep Voice 3: Scaling Text-to-Speech with Convolutional Sequence\n Learning
2017 · arXiv (Cornell University)
We present Deep Voice 3, a fully-convolutional attention-based neural\ntext-to-speech (TTS) system. Deep Voice 3 matches state-of-the-art neural\nspeech synthesis systems in naturalness while training ten times faster. We\nscale Deep Voice 3 to data set sizes unprecedented …
-
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
2019 · arXiv (Cornell University)
Transfer learning, where a model is first pre-trained on a data-rich task before being fine-tuned on a downstream task, has emerged as a powerful technique in natural language processing (NLP). The effectiveness of transfer learning …
-
ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models
2022 · Transactions of the Association for Computational Linguistics
Abstract Most widely used pre-trained language models operate on sequences of tokens corresponding to word or subword units. By comparison, token-free models that operate directly on raw text (bytes or characters) have many benefits: They …
-
PaLM: Scaling Language Modeling with Pathways
2022 · arXiv (Cornell University)
Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to …
-
Scaling Instruction-Finetuned Language Models
2022 · arXiv (Cornell University)
Finetuning language models on a collection of datasets phrased as instructions has been shown to improve model performance and generalization to unseen tasks. In this paper we explore instruction finetuning with a particular focus on …
-
Llama 2: Open Foundation and Fine-Tuned Chat Models
2023 · arXiv (Cornell University)
In this work, we develop and release Llama 2, a collection of pretrained and fine-tuned large language models (LLMs) ranging in scale from 7 billion to 70 billion parameters. Our fine-tuned LLMs, called Llama 2-Chat, …