Noam Shazeer
9 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
End-to-end text-dependent speaker verification
2016
In this paper we present a data-driven, integrated approach to speaker verification, which maps a test utterance and a few reference utterances directly to a single score for verification and jointly optimizes the system's components …
-
Exploring the Limits of Language Modeling
2016 · arXiv (Cornell University)
In this work we explore recent advances in Recurrent Neural Networks for large scale Language Modeling, a task central to language understanding. We extend current models to deal with two key challenges present in this …
-
Attention Is All You Need
2025
The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a …
-
Fast Decoding in Sequence Models using Discrete Latent Variables
2018 · arXiv (Cornell University)
Autoregressive sequence models based on deep neural networks, such as RNNs, Wavenet and the Transformer attain state-of-the-art results on many tasks. However, they are difficult to parallelize and are thus slow at processing long sequences. …
-
Corpora Generation for Grammatical Error Correction
2019
Jared Lichtarge, Chris Alberti, Shankar Kumar, Noam Shazeer, Niki Parmar, Simon Tong. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and …
-
Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
2015 · arXiv (Cornell University)
Recurrent Neural Networks can be trained to produce sequences of tokens given some input, as exemplified by recent results in machine translation and image captioning. The current approach to training them consists of maximizing the …
-
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
2019 · arXiv (Cornell University)
Transfer learning, where a model is first pre-trained on a data-rich task before being fine-tuned on a downstream task, has emerged as a powerful technique in natural language processing (NLP). The effectiveness of transfer learning …
-
PaLM: Scaling Language Modeling with Pathways
2022 · arXiv (Cornell University)
Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to …
-
LaMDA: Language Models for Dialog Applications
2022 · arXiv (Cornell University)
We present LaMDA: Language Models for Dialog Applications. LaMDA is a family of Transformer-based neural language models specialized for dialog, which have up to 137B parameters and are pre-trained on 1.56T words of public dialog …