Researcher profile

Luke Zettlemoyer

40 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Broad-coverage CCG Semantic Parsing with AMR

    2015

    We propose a grammar induction technique for AMR semantic parsing. While previous grammar induction techniques were designed to re-learn a new parser for each target application, the recently annotated AMR Bank provides a unique opportunity …

  2. Knowledge Guided Text Retrieval and Reading for Open Domain Question Answering

    2019 · arXiv (Cornell University)

    We introduce an approach for open-domain question answering (QA) that retrieves and reads a passage graph, where vertices are passages of text and edges represent relationships that are derived from an external knowledge base or …

  3. Muppet: Massive Multi-task Representations with Pre-Finetuning

    2021 · arXiv (Cornell University)

    We propose pre-finetuning, an additional large-scale learning stage between language model pre-training and fine-tuning. Pre-finetuning is massively multi-task learning (around 50 datasets, over 4.8 million total labeled examples), and is designed to encourage learning of …

  4. ART: Automatic multi-step reasoning and tool-use for large language models

    2023 · arXiv (Cornell University)

    Large language models (LLMs) can perform complex reasoning in few- and zero-shot settings by generating intermediate chain of thought (CoT) reasoning steps. Further, each reasoning step can rely on external tools to support computation beyond …

  5. Translate to Disambiguate: Zero-shot Multilingual Word Sense Disambiguation with Pretrained Language Models

    2023 · arXiv (Cornell University)

    Pretrained Language Models (PLMs) learn rich cross-lingual knowledge and can be finetuned to perform well on diverse tasks such as translation and multilingual word sense disambiguation (WSD). However, they often struggle at disambiguating word sense …

  6. One Embedder, Any Task: Instruction-Finetuned Text Embeddings

    2023

    Hongjin Su, Weijia Shi, Jungo Kasai, Yizhong Wang, Yushi Hu, Mari Ostendorf, Wen-tau Yih, Noah A. Smith, Luke Zettlemoyer, Tao Yu. Findings of the Association for Computational Linguistics: ACL 2023. 2023.

  7. 2 OLMo 2 Furious

    2024 · arXiv (Cornell University)

    We present OLMo 2, the next generation of our fully open language models. OLMo 2 includes a family of dense autoregressive language models at 7B, 13B and 32B scales with fully released artifacts -- model …

  8. LSTM CCG Parsing

    2016

    We demonstrate that a state-of-the-art parser can be built using only a lexical tagging model and a deterministic grammar, with no explicit model of bi-lexical dependencies.Instead, all dependencies are implicitly encoded in an LSTM supertagger …

  9. Summarizing Source Code using a Neural Attention Model

    2016

    High quality source code is often paired with high level summaries of the computation it performs, for example in code documentation or in descriptions posted in online forums. Such summaries are extremely useful for applications …

  10. Globally Coherent Text Generation with Neural Checklist Models

    2016

    and what still needs to be said -especially when constructing long texts. We present the neural checklist model, a recurrent neural network that models global coherence by storing and updating an agenda of text strings …

  11. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

    2017

    We present TriviaQA, a challenging reading comprehension dataset containing over 650K question-answer-evidence triples. TriviaQA includes 95K questionanswer pairs authored by trivia enthusiasts and independently gathered evidence documents, six per question on average, that provide high …

  12. Deep Semantic Role Labeling: What Works and What’s Next

    2017

    We introduce a new deep learning model for semantic role labeling (SRL) that significantly improves the state of the art, along with detailed analyses to reveal its strengths and limitations. We use a deep highway …

  13. Deep Contextualized Word Representations

    2018

    Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, Luke Zettlemoyer. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume …

  14. AllenNLP: A Deep Semantic Natural Language Processing Platform

    2018

    Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson F. Liu, Matthew Peters, Michael Schmitz, Luke Zettlemoyer. Proceedings of Workshop for NLP Open Source Software (NLP-OSS). 2018.

  15. Adversarial Example Generation with Syntactically Controlled Paraphrase Networks

    2018

    Mohit Iyyer, John Wieting, Kevin Gimpel, Luke Zettlemoyer. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.

  16. Supervised Open Information Extraction

    2018

    Gabriel Stanovsky, Julian Michael, Luke Zettlemoyer, Ido Dagan. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.

  17. QuAC: Question Answering in Context

    2018

    We present QuAC, a dataset for Question Answering in Context that contains 14K information-seeking QA dialogs (100K questions in total). The dialogs involve two crowd workers: (1) a student who poses a sequence of freeform …

  18. Dissecting Contextual Word Embeddings: Architecture and Representation

    2018

    Contextual word representations derived from pre-trained bidirectional language models (biLMs) have recently been shown to provide significant improvements to the state of the art for a wide range of NLP tasks. However, many questions remain …

  19. Transformers with convolutional context for ASR

    2019 · arXiv (Cornell University)

    The recent success of transformer networks for neural machine translation and other NLP tasks has led to a surge in research work trying to apply it for speech recognition. Recent efforts studied key research questions …

  20. Multi-hop Reading Comprehension through Question Decomposition and Rescoring

    2019

    Multi-hop Reading Comprehension (RC) requires reasoning and aggregation across several paragraphs. We propose a system for multi-hop RC that decomposes a compositional question into simpler sub-questions that can be answered by off-the-shelf single-hop RC models. …

  21. Compositional Questions Do Not Necessitate Multi-hop Reasoning

    2019

    Multi-hop reading comprehension (RC) questions are challenging because they require reading and reasoning over multiple paragraphs. We argue that it can be difficult to construct large multi-hop RC datasets. For example, even highly compositional questions …

  22. SpanBERT: Improving Pre-training by Representing and Predicting Spans

    2020 · Transactions of the Association for Computational Linguistics

    We present SpanBERT, a pre-training method that is designed to better represent and predict spans of text. Our approach extends BERT by (1) masking contiguous random spans, rather than random tokens, and (2) training the …

  23. Ultra-Fine Entity Typing

    2018

    We introduce a new entity typing task: given a sentence with an entity mention, the goal is to predict a set of free-form phrases (e.g. skyscraper, songwriter, or criminal) that describe appropriate types for the …

  24. Jointly Predicting Predicates and Arguments in Neural Semantic Role Labeling

    2018

    Recent BIO-tagging-based neural semantic role labeling models are very high performing, but assume gold predicates as part of the input and cannot incorporate span-level features. We propose an endto-end approach for jointly predicting all predicates, …

  25. HISTORIAE, History of Socio-Cultural Transformation as Linguistic Data Science. A Humanities Use Case

    2019 · DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)

    Given a combinatorial optimisation problem, there are typically multiple ways of modelling it for presentation to an automated solver. Choosing the right combination of model and target solver can have a significant impact on the …

  26. Don’t Take the Easy Way Out: Ensemble Based Methods for Avoiding Known Dataset Biases

    2019

    Christopher Clark, Mark Yatskar, Luke Zettlemoyer. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.

  27. A Discrete Hard EM Approach for Weakly Supervised Question Answering

    2019

    Sewon Min, Danqi Chen, Hannaneh Hajishirzi, Luke Zettlemoyer. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.

  28. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension

    2020

    Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, Luke Zettlemoyer. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.

  29. Unsupervised Cross-lingual Representation Learning at Scale

    2020

    Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, Veselin Stoyanov. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.

  30. Multilingual Denoising Pre-training for Neural Machine Translation

    2020 · Transactions of the Association for Computational Linguistics

    This paper demonstrates that multilingual denoising pre-training produces significant performance gains across a wide variety of machine translation (MT) tasks. We present mBART—a sequence-to-sequence denoising auto-encoder pre-trained on large-scale monolingual corpora in many languages using …

  31. Emerging Cross-lingual Structure in Pretrained Language Models

    2020

    We study the problem of multilingual masked language modeling, i.e. the training of a single model on concatenated text from multiple languages, and present a detailed study of several factors that influence why these models …

  32. AmbigQA: Answering Ambiguous Open-domain Questions

    2020

    Ambiguity is inherent to open-domain question answering; especially when exploring new topics, it can be difficult to ask questions that have a single, unambiguous answer. In this paper, we introduce AMBIGQA, a new open-domain question …

  33. Analysing Off-The-Shelf Options for Question Answering with Portuguese FAQs

    2022 · DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)

    Following the current interest in developing automatic question answering systems, we analyse alternative approaches for finding suitable answers from a list of Frequently Asked Questions (FAQs), in Portuguese. These rely on different technologies, some more …

  34. Noisy Channel Language Model Prompting for Few-Shot Text Classification

    2021 · arXiv (Cornell University)

    We introduce a noisy channel approach for language model prompting in few-shot text classification. Instead of computing the likelihood of the label given the input (referred as direct models), channel models compute the conditional probability …

  35. OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization

    2022 · arXiv (Cornell University)

    Recent work has shown that fine-tuning large pre-trained language models on a collection of tasks described via instructions, a.k.a. instruction-tuning, improves their zero and few-shot generalization to unseen tasks. However, there is a limited understanding …

  36. MetaICL: Learning to Learn In Context

    2021 · arXiv (Cornell University)

    We introduce MetaICL (Meta-training for In-Context Learning), a new meta-training framework for few-shot learning where a pretrained language model is tuned to do in-context learning on a large set of training tasks. This meta-training enables …

  37. Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?

    2022

    Large language models (LMs) are able to in-context learn—perform a new task via inference alone by conditioning on a few input-label pairs (demonstrations) and making predictions for new inputs. However, there has been little understanding …

  38. UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language Models

    2022

    Tianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao, Dragomir …

  39. FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation

    2023

    Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, Hannaneh Hajishirzi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.

  40. Demystifying Prompts in Language Models via Perplexity Estimation

    2023

    Language models can be prompted to perform a wide variety of tasks with zero- and few-shot in-context learning. However, performance varies significantly with the choice of prompt, and we do not yet understand why this …