ملف الباحث
Yangyang Zhao
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Automatic Curriculum Learning With Over-repetition Penalty for Dialogue Policy Learning
2021 · Proceedings of the AAAI Conference on Artificial Intelligence
Dialogue policy learning based on reinforcement learning is difficult to be applied to real users to train dialogue agents from scratch because of the high cost. User simulators, which choose random user goals for the …