Rewon Child
5 papers in the PaperMetrix corpus
Papers by this author
-
Active Learning for Speech Recognition: the Power of Gradients
2016 · arXiv (Cornell University)
In training speech recognition systems, labeling audio clips can be expensive, and not all data is equally valuable. Active learning aims to label only the most informative samples to reduce cost. For speech recognition, confidence …
-
Scaling Laws for Neural Language Models
2020 · arXiv (Cornell University)
This paper develops a transport-validity theory for agentic AI interventions that are first screened on small systems and later considered for frontier-scale deployment. Rather than predicting absolute frontier performance, it asks when a comparative gain …
-
Language Models are Few-Shot Learners
2020 · arXiv (Cornell University)
Recent work has demonstrated substantial gains on many NLP tasks and benchmarks by pre-training on a large corpus of text followed by fine-tuning on a specific task. While typically task-agnostic in architecture, this method still …
-
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
2022 · arXiv (Cornell University)
Pretrained general-purpose language models can achieve state-of-the-art accuracies in various natural language processing domains by adapting to downstream tasks via zero-shot, few-shot and fine-tuning techniques. Because of their success, the size of these models has …
-
PaLM: Scaling Language Modeling with Pathways
2022 · arXiv (Cornell University)
Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to …