ملف الباحث
Shikib Mehri
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Data-Efficient Alignment of Large Language Models with Human Feedback Through Natural Language
2023 · arXiv (Cornell University)
Learning from human feedback is a prominent technique to align the output of large language models (LLMs) with human expectations. Reinforcement learning from human feedback (RLHF) leverages human preference signals that are in the form …