ملف الباحث

Chelsea Finn

11 ورقة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Deep Reinforcement Learning for Vision-Based Robotic Grasping: A Simulated Comparative Evaluation of Off-Policy Methods

    2018

    In this paper, we explore deep reinforcement learning algorithms for vision-based robotic grasping. Model-free deep reinforcement learning (RL) has been successfully applied to a range of challenging environments, but the proliferation of algorithms makes it …

  2. Model-Based Reinforcement Learning for Atari

    2019 · arXiv (Cornell University)

    Model-free reinforcement learning (RL) can be used to learn effective policies for complex tasks, such as Atari games, even from image observations. However, this typically requires very large amounts of interaction -- substantially more, in …

  3. Unsupervised Curricula for Visual Meta-Reinforcement Learning

    2019 · arXiv (Cornell University)

    In principle, meta-reinforcement learning algorithms leverage experience across many tasks to learn fast reinforcement learning (RL) strategies that transfer to similar tasks. However, current meta-RL approaches rely on manually-defined distributions of training tasks, and hand-crafting …

  4. One Solution is Not All You Need: Few-Shot Extrapolation via Structured MaxEnt RL

    2020 · arXiv (Cornell University)

    While reinforcement learning algorithms can learn effective policies for complex tasks, these policies are often brittle to even minor task variations, especially when variations are not explicitly provided during training. One natural approach to this …

  5. Model-Based Visual Planning with Self-Supervised Functional Distances

    2020 · arXiv (Cornell University)

    A generalist robot must be able to complete a variety of tasks in its environment. One appealing way to specify each task is in terms of a goal observation. However, learning goal-reaching policies with reinforcement …

  6. Information is Power: Intrinsic Control via Information Capture

    2021 · arXiv (Cornell University)

    Humans and animals explore their environment and acquire useful skills even in the absence of clear goals, exhibiting intrinsic motivation. The study of intrinsic motivation in artificial agents is concerned with the following question: what …

  7. Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning

    2023 · arXiv (Cornell University)

    A compelling use case of offline reinforcement learning (RL) is to obtain a policy initialization from existing datasets followed by fast online fine-tuning with limited interaction. However, existing offline RL methods tend to behave poorly …

  8. Universal Neural Functionals

    2024 · arXiv (Cornell University)

    A challenging problem in many modern machine learning tasks is to process weight-space features, i.e., to transform or extract information from the weights and gradients of a neural network. Recent works have developed promising weight-space …

  9. A Critical Evaluation of AI Feedback for Aligning Large Language Models

    2024 · arXiv (Cornell University)

    Reinforcement learning with AI feedback (RLAIF) is a popular paradigm for improving the instruction-following abilities of powerful pre-trained language models. RLAIF first performs supervised fine-tuning (SFT) using demonstrations from a teacher model and then further …

  10. D5RL: Diverse Datasets for Data-Driven Deep Reinforcement Learning

    2024 · arXiv (Cornell University)

    Offline reinforcement learning algorithms hold the promise of enabling data-driven RL methods that do not require costly or dangerous real-world exploration and benefit from large pre-collected datasets. This in turn can facilitate real-world applications, as …

  11. Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

    2022 · arXiv (Cornell University)

    Large language models can encode a wealth of semantic knowledge about the world. Such knowledge could be extremely useful to robots aiming to act upon high-level, temporally extended instructions expressed in natural language. However, a …