Owain Evans
4 papers in the PaperMetrix corpus
Papers by this author
-
TruthfulQA: Measuring How Models Mimic Human Falsehoods
2021 · arXiv (Cornell University)
We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions …
-
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
2024 · arXiv (Cornell University)
One way to address safety risks from large language models (LLMs) is to censor dangerous knowledge from their training data. While this removes the explicit information, implicit information can remain scattered across various training documents. …
-
Looking Inward: Language Models Can Learn About Themselves by Introspection
2024 · arXiv (Cornell University)
Humans acquire knowledge by observing the external world, but also by introspection. Introspection gives a person privileged access to their current state of mind (e.g., thoughts and feelings) that is not accessible to external observers. …
-
Training large language models on narrow tasks can lead to broad misalignment
2026 · Nature
Abstract The widespread adoption of large language models (LLMs) raises important questions about their safety and alignment 1 . Previous safety research has largely focused on isolated undesirable behaviours, such as reinforcing harmful stereotypes or …