Yejin Choi
28 papers in the PaperMetrix corpus
Papers by this author
-
Scarecrow: A Framework for Scrutinizing Machine Text.
2021 · arXiv (Cornell University)
Modern neural text generation systems can produce remarkably fluent and grammatical texts. While earlier language models suffered from repetition and syntactic errors, the errors made by contemporary models are often semantic, narrative, or discourse failures. …
-
Symbolic Knowledge Distillation: from General Language Models to Commonsense Models
2022 · Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
Peter West, Chandra Bhagavatula, Jack Hessel, Jena Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, Yejin Choi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: …
-
Aligning to Social Norms and Values in Interactive Narratives
2022 · arXiv (Cornell University)
We focus on creating agents that act in alignment with socially beneficial norms and values in interactive narratives or text-based games -- environments wherein an agent perceives and interacts with a world through natural language. …
-
Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
2022 · arXiv (Cornell University)
We tackle the problem of aligning pre-trained large language models (LMs) with human preferences. If we view text generation as a sequential decision-making problem, reinforcement learning (RL) appears to be a natural conceptual framework. However, …
-
PlaSma: Making Small Language Models Better Procedural Knowledge Models for (Counterfactual) Planning
2023 · arXiv (Cornell University)
Procedural planning, which entails decomposing a high-level goal into a sequence of temporally ordered steps, is an important yet intricate task for machines. It involves integrating common-sense knowledge to reason about complex and often contextualized …
-
Symbolic Chain-of-Thought Distillation: Small Models Can Also "Think" Step-by-Step
2023 · arXiv (Cornell University)
Chain-of-thought prompting (e.g., "Let's think step-by-step") primes large language models to verbalize rationalization for their predictions. While chain-of-thought can lead to dramatic performance gains, benefits appear to emerge only for sufficiently large models (beyond 50B …
-
STEER: Unified Style Transfer with Expert Reinforcement
2023 · arXiv (Cornell University)
While text style transfer has many applications across natural language processing, the core premise of transferring from a single source style is unrealistic in a real-world setting. In this work, we focus on arbitrary style …
-
MacGyver: Are Large Language Models Creative Problem Solvers?
2024
Yufei Tian, Abhilasha Ravichander, Lianhui Qin, Ronan Le Bras, Raja Marjieh, Nanyun Peng, Yejin Choi, Thomas Griffiths, Faeze Brahman. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: …
-
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
2025 · arXiv (Cornell University)
We investigate the logical reasoning capabilities of large language models (LLMs) and their scalability in complex non-monotonic reasoning. To this end, we introduce ZebraLogic, a comprehensive evaluation framework for assessing LLM reasoning performance on logic …
-
OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens
2025 · arXiv (Cornell University)
We present OLMoTrace, the first system that traces the outputs of language models back to their full, multi-trillion-token training data in real time. OLMoTrace finds and shows verbatim matches between segments of language model output …
-
Globally Coherent Text Generation with Neural Checklist Models
2016
and what still needs to be said -especially when constructing long texts. We present the neural checklist model, a recurrent neural network that models global coherence by storing and updating an agenda of text strings …
-
Generating Topical Poetry
2016
We describe Hafez, a program that generates any number of distinct poems on a usersupplied topic. Poems obey rhythmic and rhyme constraints. We describe the poetrygeneration algorithm, give experimental data concerning its parameters, and show …
-
Deep Communicating Agents for Abstractive Summarization
2018
Asli Celikyilmaz, Antoine Bosselut, Xiaodong He, Yejin Choi. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
-
SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference
2018
Given a partial description like "she opened the hood of the car," humans can reason about the situation and anticipate what might come next ("then, she examined the engine"). In this paper, we introduce the …
-
QuAC: Question Answering in Context
2018
We present QuAC, a dataset for Question Answering in Context that contains 14K information-seeking QA dialogs (100K questions in total). The dialogs involve two crowd workers: (1) a student who poses a sequence of freeform …
-
DREAM: A Challenge Data Set and Models for Dialogue-Based Reading Comprehension
2019 · Transactions of the Association for Computational Linguistics
We present DREAM, the first dialogue-based multiple-choice reading comprehension data set. Collected from English as a Foreign Language examinations designed by human experts to evaluate the comprehension level of Chinese learners of English, our data …
-
Social IQa: Commonsense Reasoning about Social Interactions
2019
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, Yejin Choi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
-
The Curious Case of Neural Text Degeneration
2019 · arXiv (Cornell University)
Despite considerable advancements with deep neural language models, the enigma of neural text degeneration persists when these models are tested as text generators. The counter-intuitive empirical observation is that even though the use of likelihood …
-
HellaSwag: Can a Machine Really Finish Your Sentence?
2019
introduced a new task of commonsense natural language inference: given an event description such as "A woman sits at a piano," a machine must select the most likely followup: "She sets her fingers on the …
-
COMET: Commonsense Transformers for Automatic Knowledge Graph Construction
2019
We present the first comprehensive study on automatic knowledge base construction for two prevalent commonsense knowledge graphs: ATOMIC Contrary to many conventional KBs that store knowledge with canonical templates, commonsense KBs only store loosely structured …
-
Ultra-Fine Entity Typing
2018
We introduce a new entity typing task: given a sentence with an entity mention, the goal is to predict a set of free-form phrases (e.g. skyscraper, songwriter, or criminal) that describe appropriate types for the …
-
ATOMIC: An Atlas of Machine Commonsense for If-Then Reasoning
2019
We present ATOMIC, an atlas of everyday commonsense reasoning, organized through 877k textual descriptions of inferential knowledge. Compared to existing resources that center around taxonomic knowledge, ATOMIC focuses on inferential knowledge organized as typed if-then …
-
Cosmos QA: Machine Reading Comprehension with Contextual Commonsense Reasoning
2019
Lifu Huang, Ronan Le Bras, Chandra Bhagavatula, Yejin Choi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
-
The Curious Case of Neural Text Degeneration
2020 · arXiv (Cornell University)
Despite considerable advances in neural language modeling, it remains an open question what the best decoding strategy is for text generation from a language model (e.g. to generate a story). The counter-intuitive empirical observation is …
-
WinoGrande: An Adversarial Winograd Schema Challenge at Scale
2020 · Proceedings of the AAAI Conference on Artificial Intelligence
The Winograd Schema Challenge (WSC) (Levesque, Davis, and Morgenstern 2011), a benchmark for commonsense reasoning, is a set of 273 expert-crafted pronoun resolution problems originally designed to be unsolvable for statistical models that rely on …
-
PIQA: Reasoning about Physical Commonsense in Natural Language
2020
To apply eyeshadow without a brush, should I use a cotton swab or a toothpick? Questions requiring this kind of physical commonsense pose a challenge to today's natural language understanding systems. While recent pretrained models …
-
Generative Data Augmentation for Commonsense Reasoning
2020
Yiben Yang, Chaitanya Malaviya, Jared Fernandez, Swabha Swayamdipta, Ronan Le Bras, Ji-Ping Wang, Chandra Bhagavatula, Yejin Choi, Doug Downey. Findings of the Association for Computational Linguistics: EMNLP 2020. 2020.
-
(Comet-) Atomic 2020: On Symbolic and Neural Commonsense Knowledge Graphs
2021
Recent years have brought about a renewed interest in commonsense representation and reasoning in the field of natural language understanding. The development of new commonsense knowledge graphs (CSKG) has been central to these advances as …