ملف الباحث
Yehong Zhang
ورقتان في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
COPR: Continual Human Preference Learning via Optimal Policy Regularization
2024 · arXiv (Cornell University)
Reinforcement Learning from Human Feedback (RLHF) is commonly utilized to improve the alignment of Large Language Models (LLMs) with human preferences. Given the evolving nature of human preferences, continual alignment becomes more crucial and practical …
-
PanGu-$α$: Large-scale Autoregressive Pretrained Chinese Language Models with Auto-parallel Computation
2021 · arXiv (Cornell University)
Large-scale Pretrained Language Models (PLMs) have become the new paradigm for Natural Language Processing (NLP). PLMs with hundreds of billions parameters such as GPT-3 have demonstrated strong performances on natural language understanding and generation with …