Huazheng Wang
3 papers in the PaperMetrix corpus
Papers by this author
-
Provably Efficient Representation Learning with Tractable Planning in Low-Rank POMDP
2023 · arXiv (Cornell University)
In this paper, we study representation learning in partially observable Markov Decision Processes (POMDPs), where the agent learns a decoder function that maps a series of high-dimensional raw observations to a compact representation and uses …
-
RA-PbRL: Provably Efficient Risk-Aware Preference-Based Reinforcement Learning
2024 · arXiv (Cornell University)
Reinforcement Learning from Human Feedback (RLHF) has recently surged in popularity, particularly for aligning large language models and other AI systems with human intentions. At its core, RLHF can be viewed as a specialized instance …
-
Factorization Bandits for Interactive Recommendation
2017 · Proceedings of the AAAI Conference on Artificial Intelligence
We perform online interactive recommendation via a factorization-based bandit algorithm. Low-rank matrix completion is performed over an incrementally constructed user-item preference matrix, where an upper confidence bound based item selection strategy is developed to balance …