Jiahui Li
3 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Shapley Counterfactual Credits for Multi-Agent Reinforcement Learning
2021
Centralized Training with Decentralized Execution (CTDE) has been a popular paradigm in cooperative Multi-Agent Reinforcement Learning (MARL) settings and is widely used in many real applications. One of the major challenges in the training process …
-
CBDA: Chaos-based binary dragonfly algorithm for evolutionary feature selection
2024 · Intelligent Data Analysis
The goal of feature selection in machine learning is to simultaneously maintain more classification accuracy, while reducing lager amount of attributes. In this paper, we firstly design a fitness function that achieves both objectives jointly. …
-
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution
2024 · arXiv (Cornell University)
Reinforcement learning from human feedback (RLHF) offers a promising approach to aligning large language models (LLMs) with human preferences. Typically, a reward model is trained or supplied to act as a proxy for humans in …