Lin Gui
3 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
A Question Answering Approach to Emotion Cause Extraction
2017 · arXiv (Cornell University)
Emotion cause extraction aims to identify the reasons behind a certain emotion expressed in text. It is a much more difficult task compared to emotion classification. Inspired by recent advances in using deep memory networks …
-
COPR: Continual Human Preference Learning via Optimal Policy Regularization
2024 · arXiv (Cornell University)
Reinforcement Learning from Human Feedback (RLHF) is commonly utilized to improve the alignment of Large Language Models (LLMs) with human preferences. Given the evolving nature of human preferences, continual alignment becomes more crucial and practical …
-
Let it Calm: Exploratory Annealed Decoding for Verifiable Reinforcement Learning
2025 · arXiv (Cornell University)
Reinforcement learning with verifiable rewards (RLVR) is a powerful paradigm for enhancing the reasoning capabilities of large language models (LLMs), yet its success hinges on effective exploration. An ideal exploration strategy must navigate two fundamental …