Researcher profile

Owain Evans

4 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. TruthfulQA: Measuring How Models Mimic Human Falsehoods

    2021 · arXiv (Cornell University)

    We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions …

  2. Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data

    2024 · arXiv (Cornell University)

    One way to address safety risks from large language models (LLMs) is to censor dangerous knowledge from their training data. While this removes the explicit information, implicit information can remain scattered across various training documents. …

  3. Looking Inward: Language Models Can Learn About Themselves by Introspection

    2024 · arXiv (Cornell University)

    Humans acquire knowledge by observing the external world, but also by introspection. Introspection gives a person privileged access to their current state of mind (e.g., thoughts and feelings) that is not accessible to external observers. …

  4. Training large language models on narrow tasks can lead to broad misalignment

    2026 · Nature

    Abstract The widespread adoption of large language models (LLMs) raises important questions about their safety and alignment 1 . Previous safety research has largely focused on isolated undesirable behaviours, such as reinforcing harmful stereotypes or …