Researcher profile
Zixuan Zhang
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
Robust Reinforcement Learning from Corrupted Human Feedback
2024 · arXiv (Cornell University)
Reinforcement learning from human feedback (RLHF) provides a principled framework for aligning AI systems with human preference data. For various reasons, e.g., personal bias, context ambiguity, lack of training, etc, human annotators may give incorrect …