Denny Zhou
9 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Doubly Sparse: Sparse Mixture of Sparse Experts for Efficient Softmax Inference
2019 · arXiv (Cornell University)
Computations for the softmax function are significantly expensive when the number of output classes is large. In this paper, we present a novel softmax inference speedup method, Doubly Sparse Softmax (DS-Softmax), that leverages sparse mixture …
-
Emergent Abilities of Large Language Models
2022 · arXiv (Cornell University)
Scaling up language models has been shown to predictably improve performance and sample efficiency on a wide range of downstream tasks. This paper instead discusses an unpredictable phenomenon that we refer to as emergent abilities …
-
Symbol tuning improves in-context learning in language models
2023 · arXiv (Cornell University)
We present symbol tuning - finetuning language models on in-context input-label pairs where natural language labels (e.g., "positive/negative sentiment") are replaced with arbitrary symbols (e.g., "foo/bar"). Symbol tuning leverages the intuition that when a model …
-
MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices
2020
Natural Language Processing (NLP) has recently achieved great success by using huge pre-trained models with hundreds of millions of parameters. However, these models suffer from heavy model sizes and high latency such that they cannot …
-
BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
2022 · arXiv (Cornell University)
AbstractThere is a failure mode in large language models that we do not have a good name for, and thatwe therefore tend not to treat seriously enough. It is not hallucination — the model is …
-
PaLM: Scaling Language Modeling with Pathways
2022 · arXiv (Cornell University)
Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to …
-
Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
2022 · arXiv (Cornell University)
Chain-of-thought prompting has demonstrated remarkable performance on various natural language reasoning tasks. However, it tends to perform poorly on tasks which requires solving problems harder than the exemplars shown in the prompts. To overcome this …
-
UL2: Unifying Language Learning Paradigms
2022 · arXiv (Cornell University)
Existing pre-trained models are generally geared towards a particular class of problems. To date, there seems to be still no consensus on what the right architecture and pre-training setup should be. This paper presents a …
-
Scaling Instruction-Finetuned Language Models
2022 · arXiv (Cornell University)
Finetuning language models on a collection of datasets phrased as instructions has been shown to improve model performance and generalization to unseen tasks. In this paper we explore instruction finetuning with a particular focus on …