ملف الباحث
Sung‐Jin Lee
ورقتان في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Data-Efficient Alignment of Large Language Models with Human Feedback Through Natural Language
2023 · arXiv (Cornell University)
Learning from human feedback is a prominent technique to align the output of large language models (LLMs) with human expectations. Reinforcement learning from human feedback (RLHF) leverages human preference signals that are in the form …
-
Jointly Optimizing Diversity and Relevance in Neural Response Generation
2019
Xiang Gao, Sungjin Lee, Yizhe Zhang, Chris Brockett, Michel Galley, Jianfeng Gao, Bill Dolan. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 …