ملف الباحث

Hanpeng Hu

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Optimizing RLHF Training for Large Language Models with Stage Fusion

    2024 · arXiv (Cornell University)

    We present RLHFuse, an efficient training system with stage fusion for Reinforcement Learning from Human Feedback (RLHF). Due to the intrinsic nature of RLHF training, i.e., the data skewness in the generation stage and the …