Researcher profile

Leo Gao

3 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Scaling Laws for Reward Model Overoptimization

    2022 · arXiv (Cornell University)

    In reinforcement learning from human feedback, it is common to optimize against a reward model trained to predict human preferences. Because the reward model is an imperfect proxy, optimizing its value too much can hinder …

  2. The Pile: An 800GB Dataset of Diverse Text for Language Modeling

    2020 · arXiv (Cornell University)

    Recent work has demonstrated that increased training dataset diversity improves general cross-domain knowledge and downstream generalization capability for large-scale language models. With this in mind, we present \textit{the Pile}: an 825 GiB English text corpus …

  3. Multitask Prompted Training Enables Zero-Shot Task Generalization

    2021 · arXiv (Cornell University)

    Large language models have recently been shown to attain reasonable zero-shot generalization on a diverse set of tasks (Brown et al., 2020). It has been hypothesized that this is a consequence of implicit multitask learning …