Researcher profile

Tyna Eloundou

1 paper in the PaperMetrix corpus

Publications

Papers by this author

  1. SEAL: Systematic Error Analysis for Value ALignment

    2025 · Proceedings of the AAAI Conference on Artificial Intelligence

    Reinforcement Learning from Human Feedback (RLHF) aligns language models (LMs) with human values by training reward models (RMs) on binary preferences and using these RMs to fine-tune the base models. Despite its importance, the internal …