Ofir Nachum
7 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Deep Reinforcement Learning for Vision-Based Robotic Grasping: A Simulated Comparative Evaluation of Off-Policy Methods
2018
In this paper, we explore deep reinforcement learning algorithms for vision-based robotic grasping. Model-free deep reinforcement learning (RL) has been successfully applied to a range of challenging environments, but the proliferation of algorithms makes it …
-
RL Unplugged: Benchmarks for Offline Reinforcement Learning.
2020 · arXiv (Cornell University)
Offline methods for reinforcement learning have a potential to help bridge the gap between reinforcement learning research and real-world applications. They make it possible to learn policies from offline datasets, thus overcoming concerns associated with …
-
Identifying and Correcting Label Bias in Machine Learning
2019 · International Conference on Artificial Intelligence and Statistics
The present disclosure is directed to systems and methods for identifying and correcting label bias in machine learning via intelligent re-weighting of training examples. In particular, aspects of the present disclosure leverage a problem formulation …
-
OPAL: Offline Primitive Discovery for Accelerating Offline Reinforcement Learning
2020 · arXiv (Cornell University)
Reinforcement learning (RL) has achieved impressive performance in a variety of online settings in which an agent's ability to query the environment for transitions and rewards is effectively unlimited. However, in many practical applications, the …
-
Policy Information Capacity: Information-Theoretic Measure for Task Complexity in Deep Reinforcement Learning
2021 · arXiv (Cornell University)
Progress in deep reinforcement learning (RL) research is largely enabled by benchmark task environments. However, analyzing the nature of those environments is often overlooked. In particular, we still do not have agreeable ways to measure …
-
Model Selection in Batch Policy Optimization
2021 · arXiv (Cornell University)
We study the problem of model selection in batch policy optimization: given a fixed, partial-feedback dataset and $M$ model classes, learn a policy with performance that is competitive with the policy derived from the best …
-
Multi-Game Decision Transformers
2022 · arXiv (Cornell University)
A longstanding goal of the field of AI is a method for learning a highly capable, generalist agent from diverse experience. In the subfields of vision and language, this was largely achieved by scaling up …