ملف الباحث

Y. X. Zhao

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. RA-PbRL: Provably Efficient Risk-Aware Preference-Based Reinforcement Learning

    2024 · arXiv (Cornell University)

    Reinforcement Learning from Human Feedback (RLHF) has recently surged in popularity, particularly for aligning large language models and other AI systems with human intentions. At its core, RLHF can be viewed as a specialized instance …