Researcher profile
Pulkit Agrawal
2 papers in the PaperMetrix corpus
Publications
Papers by this author
-
OPAL: Offline Primitive Discovery for Accelerating Offline Reinforcement Learning
2020 · arXiv (Cornell University)
Reinforcement learning (RL) has achieved impressive performance in a variety of online settings in which an agent's ability to query the environment for transitions and rewards is effectively unlimited. However, in many practical applications, the …
-
Value Augmented Sampling for Language Model Alignment and Personalization
2024 · ArXiv.org
Aligning Large Language Models (LLMs) to cater to different human preferences, learning new skills, and unlearning harmful behavior is an important problem. Search-based methods, such as Best-of-N or Monte-Carlo Tree Search, are performant, but impractical …