Researcher profile
Greg Leppert
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
SEAL: Systematic Error Analysis for Value ALignment
2025 · Proceedings of the AAAI Conference on Artificial Intelligence
Reinforcement Learning from Human Feedback (RLHF) aligns language models (LMs) with human values by training reward models (RMs) on binary preferences and using these RMs to fine-tune the base models. Despite its importance, the internal …