ملف الباحث
Feng-Lin Li
ورقتان في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Don't Forget Your Reward Values: Language Model Alignment via Value-based Calibration
2024 · arXiv (Cornell University)
While Reinforcement Learning from Human Feedback (RLHF) significantly enhances the generation quality of Large Language Models (LLMs), recent studies have raised concerns regarding the complexity and instability associated with the Proximal Policy Optimization (PPO) algorithm, …
-
AliMe Chat: A Sequence to Sequence and Rerank based Chatbot Engine
2017
Minghui Qiu, Feng-Lin Li, Siyu Wang, Xing Gao, Yan Chen, Weipeng Zhao, Haiqing Chen, Jun Huang, Wei Chu. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 2017.