Subhajit Chaudhury
3 papers in the PaperMetrix corpus
Papers by this author
-
Internal Model from Observations for Reward Shaping
2018 · arXiv (Cornell University)
Reinforcement learning methods require careful design involving a reward function to obtain the desired action policy for a given task. In the absence of hand-crafted reward functions, prior work on the topic has proposed several …
-
Constrained Exploration and Recovery from Experience Shaping
2018 · arXiv (Cornell University)
We consider the problem of reinforcement learning under safety requirements, in which an agent is trained to complete a given task, typically formalized as the maximization of a reward signal over time, while concurrently avoiding …
-
Scalable Learning of Latent Language Structure With Logical Offline Cycle Consistency
2023 · arXiv (Cornell University)
We introduce Logical Offline Cycle Consistency Optimization (LOCCO), a scalable, semi-supervised method for training a neural semantic parser. Conceptually, LOCCO can be viewed as a form of self-learning where the semantic parser being trained is …