Bin Liang
3 papers in the PaperMetrix corpus
Papers by this author
-
Distributed Deep Reinforcement Learning: A Survey and A Multi-Player Multi-Agent Learning Toolbox
2022 · arXiv (Cornell University)
With the breakthrough of AlphaGo, deep reinforcement learning becomes a recognized technique for solving sequential decision-making problems. Despite its reputation, data inefficiency caused by its trial and error learning mechanism makes deep reinforcement learning hard …
-
MMSD2.0: Towards a Reliable Multi-modal Sarcasm Detection System
2023
Multi-modal sarcasm detection has attracted much recent attention. Nevertheless, the existing benchmark (MMSD) has some shortcomings that hinder the development of reliable multi-modal sarcasm detection system: (1) There are some spurious cues in MMSD, leading …
-
COPR: Continual Human Preference Learning via Optimal Policy Regularization
2024 · arXiv (Cornell University)
Reinforcement Learning from Human Feedback (RLHF) is commonly utilized to improve the alignment of Large Language Models (LLMs) with human preferences. Given the evolving nature of human preferences, continual alignment becomes more crucial and practical …