Researcher profile

Jacob Hilton

3 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Scaling Laws for Reward Model Overoptimization

    2022 · arXiv (Cornell University)

    In reinforcement learning from human feedback, it is common to optimize against a reward model trained to predict human preferences. Because the reward model is an imperfect proxy, optimizing its value too much can hinder …

  2. TruthfulQA: Measuring How Models Mimic Human Falsehoods

    2021 · arXiv (Cornell University)

    We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions …

  3. Training language models to follow instructions with human feedback

    2022 · arXiv (Cornell University)

    Making language models bigger does not inherently make them better at following a user's intent. For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user. In …