Bryan Catanzaro
5 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
DSD: Regularizing Deep Neural Networks with Dense-Sparse-Dense Training Flow.
2016 · arXiv (Cornell University)
Modern deep neural networks have a large number of parameters, making them very powerful machine learning systems. A critical issue for training such large networks on large-scale data-sets is to prevent overfitting while at the …
-
Large Scale Multi-Actor Generative Dialog Modeling
2020
Non-goal oriented dialog agents (i.e. chatbots) aim to produce varying and engaging conversations with a user; however, they typically exhibit either inconsistent personality across conversations or the average personality of all users. This paper addresses …
-
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
2015 · arXiv (Cornell University)
We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech--two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning …
-
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
2019 · arXiv (Cornell University)
Recent work in language modeling demonstrates that training large transformer models advances the state of the art in Natural Language Processing applications. However, very large models can be quite difficult to train due to memory …
-
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
2022 · arXiv (Cornell University)
Pretrained general-purpose language models can achieve state-of-the-art accuracies in various natural language processing domains by adapting to downstream tasks via zero-shot, few-shot and fine-tuning techniques. Because of their success, the size of these models has …