ملف الباحث

Wei, Jason

8 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Emergent Abilities of Large Language Models

    2022 · arXiv (Cornell University)

    Scaling up language models has been shown to predictably improve performance and sample efficiency on a wide range of downstream tasks. This paper instead discusses an unpredictable phenomenon that we refer to as emergent abilities …

  2. A Recipe For Arbitrary Text Style Transfer with Large Language Models

    2021 · arXiv (Cornell University)

    In this paper, we leverage large language models (LMs) to perform zero-shot text style transfer. We present a prompting method that we call augmented zero-shot learning, which frames style transfer as a sentence rewriting task …

  3. BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence

    2022 · arXiv (Cornell University)

    AbstractThere is a failure mode in large language models that we do not have a good name for, and thatwe therefore tend not to treat seriously enough. It is not hallucination — the model is …

  4. Self-Consistency Improves Chain of Thought Reasoning in Language Models

    2022 · arXiv (Cornell University)

    Chain-of-thought prompting combined with pre-trained large language models has achieved encouraging results on complex reasoning tasks. In this paper, we propose a new decoding strategy, self-consistency, to replace the naive greedy decoding used in chain-of-thought …

  5. PaLM: Scaling Language Modeling with Pathways

    2022 · arXiv (Cornell University)

    Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to …

  6. Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

    2022 · arXiv (Cornell University)

    Chain-of-thought prompting has demonstrated remarkable performance on various natural language reasoning tasks. However, it tends to perform poorly on tasks which requires solving problems harder than the exemplars shown in the prompts. To overcome this …

  7. UL2: Unifying Language Learning Paradigms

    2022 · arXiv (Cornell University)

    Existing pre-trained models are generally geared towards a particular class of problems. To date, there seems to be still no consensus on what the right architecture and pre-training setup should be. This paper presents a …

  8. Scaling Instruction-Finetuned Language Models

    2022 · arXiv (Cornell University)

    Finetuning language models on a collection of datasets phrased as instructions has been shown to improve model performance and generalization to unseen tasks. In this paper we explore instruction finetuning with a particular focus on …