Researcher profile
Shikib Mehri
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
Data-Efficient Alignment of Large Language Models with Human Feedback Through Natural Language
2023 · arXiv (Cornell University)
Learning from human feedback is a prominent technique to align the output of large language models (LLMs) with human expectations. Reinforcement learning from human feedback (RLHF) leverages human preference signals that are in the form …