ملف الباحث
Hanpeng Hu
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Optimizing RLHF Training for Large Language Models with Stage Fusion
2024 · arXiv (Cornell University)
We present RLHFuse, an efficient training system with stage fusion for Reinforcement Learning from Human Feedback (RLHF). Due to the intrinsic nature of RLHF training, i.e., the data skewness in the generation stage and the …