Researcher profile
Yuanzhao Zhai
2 papers in the PaperMetrix corpus
Publications
Papers by this author
-
Diversifying Message Aggregation in Multi-Agent Communication via Normalized Tensor Nuclear Norm Regularization
2022 · arXiv (Cornell University)
Aggregating messages is a key component for the communication of multi-agent reinforcement learning (Comm-MARL). Recently, it has witnessed the prevalence of graph attention networks (GAT) in Comm-MARL, where agents can be represented as nodes and …
-
COPR: Continual Human Preference Learning via Optimal Policy Regularization
2024 · arXiv (Cornell University)
Reinforcement Learning from Human Feedback (RLHF) is commonly utilized to improve the alignment of Large Language Models (LLMs) with human preferences. Given the evolving nature of human preferences, continual alignment becomes more crucial and practical …