Researcher profile

Dale Schuurmans

7 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. The Curse of Passive Data Collection in Batch Reinforcement Learning

    2021 · arXiv (Cornell University)

    In high stake applications, active experimentation may be considered too risky and thus data are often collected passively. While in simple cases, such as in bandits, passive and active data collection are similarly effective, the …

  2. A Simple Decentralized Cross-Entropy Method

    2022 · arXiv (Cornell University)

    Cross-Entropy Method (CEM) is commonly used for planning in model-based reinforcement learning (MBRL) where a centralized approach is typically utilized to update the sampling distribution based on only the top-$k$ operation's results on samples. In …

  3. Learning Interactive Real-World Simulators

    2023 · arXiv (Cornell University)

    Generative models trained on internet data have revolutionized how text, image, and video content can be created. Perhaps the next milestone for generative models is to simulate realistic experience in response to actions taken by …

  4. Reward Augmented Maximum Likelihood for Neural Structured Prediction

    2016 · arXiv (Cornell University)

    A key problem in structured output prediction is direct optimization of the task reward function that matters for test evaluation. This paper presents a simple and computationally efficient approach to incorporate task reward into a …

  5. BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence

    2022 · arXiv (Cornell University)

    AbstractThere is a failure mode in large language models that we do not have a good name for, and thatwe therefore tend not to treat seriously enough. It is not hallucination — the model is …

  6. Self-Consistency Improves Chain of Thought Reasoning in Language Models

    2022 · arXiv (Cornell University)

    Chain-of-thought prompting combined with pre-trained large language models has achieved encouraging results on complex reasoning tasks. In this paper, we propose a new decoding strategy, self-consistency, to replace the naive greedy decoding used in chain-of-thought …

  7. Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

    2022 · arXiv (Cornell University)

    Chain-of-thought prompting has demonstrated remarkable performance on various natural language reasoning tasks. However, it tends to perform poorly on tasks which requires solving problems harder than the exemplars shown in the prompts. To overcome this …