Researcher profile

Christopher D. Manning

32 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Universal Dependencies v1: A Multilingual Treebank Collection

    2016

    Cross-linguistically consistent annotation is necessary for sound comparative evaluation and cross-lingual learning experiments.It is also useful for multilingual system development and comparative linguistic studies.Universal Dependencies is an open community effort to create cross-linguistically consistent treebank …

  2. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

    2018 · arXiv (Cornell University)

    Existing question answering (QA) datasets fail to train QA systems to perform complex reasoning and provide explanations for answers. We introduce HotpotQA, a new dataset with 113k Wikipedia-based question-answer pairs with four key features: (1) …

  3. Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection

    2020 · Uppsala University Publications (Uppsala University)

    Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages within a dependency-based lexicalist framework. The annotation consists in a linguistically motivated word segmentation; a morphological layer comprising lemmas, …

  4. Towards Ecologically Valid Research on Language User Interfaces

    2020 · arXiv (Cornell University)

    Language User Interfaces (LUIs) could improve human-machine interaction for a wide variety of tasks, such as playing music, getting insights from databases, or instructing domestic robots. In contrast to traditional hand-crafted approaches, recent work attempts …

  5. Retrieve, Read, Rerank, then Iterate: Answering Open-Domain Questions of Varying Reasoning Steps from Text

    2020 · arXiv (Cornell University)

    We develop a unified system to answer directly from text open-domain questions that may require a varying number of retrieval steps. We employ a single multi-task transformer model to perform all the necessary subtasks -- …

  6. Pushdown Layers: Encoding Recursive Structure in Transformer Language Models

    2023 · arXiv (Cornell University)

    Recursion is a prominent feature of human language, and fundamentally challenging for self-attention due to the lack of an explicit recursive-state tracking mechanism. Consequently, Transformer language models poorly capture long-tail recursive structure and exhibit sample-inefficient …

  7. Mapping the Increasing Use of LLMs in Scientific Papers

    2024 · arXiv (Cornell University)

    Scientific publishing lays the foundation of science by disseminating research findings, fostering collaboration, encouraging reproducibility, and ensuring that scientific knowledge is accessible, verifiable, and built upon over time. Recently, there has been immense speculation about …

  8. Stronger Baselines for Retrieval-Augmented Generation with Long-Context Language Models

    2025 · arXiv (Cornell University)

    With the rise of long-context language models (LMs) capable of processing tens of thousands of tokens in a single context window, do multi-stage retrieval-augmented generation (RAG) pipelines still offer measurable benefits over simpler, single-stage approaches? …

  9. Advances in natural language processing

    2015 · Science

    Natural language processing employs computational techniques for the purpose of learning, understanding, and producing human language content. Early computational approaches to language research focused on automating the analysis of the linguistic structure of language and …

  10. A large annotated corpus for learning natural language inference

    2015 · arXiv (Cornell University)

    Understanding entailment and contradiction is fundamental to understanding natural language, and inference about entailment and contradiction is a valuable testing ground for the development of semantic representations. However, machine learning research in this area has …

  11. Effective Approaches to Attention-based Neural Machine Translation

    2015 · arXiv (Cornell University)

    An attentional mechanism has lately been used to improve neural machine translation (NMT) by selectively focusing on parts of the source sentence during translation. However, there has been little work exploring useful architectures for attention-based …

  12. Improved Semantic Representations From Tree-Structured Long Short-Term Memory Networks

    2015

    Kai Sheng Tai, Richard Socher, Christopher D. Manning. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.

  13. Bilingual Word Representations with Monolingual Quality in Mind

    2015

    Recent work in learning bilingual representations tend to tailor towards achieving good performance on bilingual tasks, most often the crosslingual document classification (CLDC) evaluation, but to the detriment of preserving clustering structures of word representations …

  14. Learning Distributed Representations for Multilingual Text Sequences

    2015

    We propose a novel approach to learning dis-tributed representations of variable-length text sequences in multiple languages simultane-ously. Unlike previous work which often de-rive representations of multi-word sequences as weighted sums of individual word vec-tors, our …

  15. Leveraging Linguistic Structure For Open Domain Information Extraction

    2015

    Gabor Angeli, Melvin Jose Johnson Premkumar, Christopher D. Manning. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). …

  16. Entity-Centric Coreference Resolution with Model Stacking

    2015

    Kevin Clark, Christopher D. Manning. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.

  17. A Fast Unified Model for Parsing and Sentence Understanding

    2016

    Samuel R. Bowman, Jon Gauthier, Abhinav Rastogi, Raghav Gupta, Christopher D. Manning, Christopher Potts. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016.

  18. Achieving Open Vocabulary Neural Machine Translation with Hybrid Word-Character Models

    2016

    Nearly all previous work on neural machine translation (NMT) has used quite restricted vocabularies, perhaps with a subsequent method to patch in unknown words. This paper presents a novel wordcharacter solution to achieving open vocabulary …

  19. Get To The Point: Summarization with Pointer-Generator Networks

    2017 · arXiv (Cornell University)

    Neural sequence-to-sequence models have provided a viable new approach for abstractive text summarization (meaning they are not restricted to simply selecting and rearranging passages from the original text). However, these models have two shortcomings: they …

  20. Stanford's Graph-based Neural Dependency Parser at the CoNLL 2017 Shared Task

    2017

    This paper describes the neural dependency parser submitted by Stanford to the CoNLL 2017 Shared Task on parsing Universal Dependencies.Our system uses relatively simple LSTM networks to produce part of speech tags and labeled dependency …

  21. CoNLL 2017 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies

    2017

    Daniel Zeman, Martin Popel, Milan Straka, Jan Hajič, Joakim Nivre, Filip Ginter, Juhani Luotolahti, Sampo Pyysalo, Slav Petrov, Martin Potthast, Francis Tyers, Elena Badmaeva, Memduh Gokirmak, Anna Nedoluzhko, Silvie Cinková, Jan Hajič jr., Jaroslava Hlaváčová, …

  22. Position-aware Attention and Supervised Data Improve Slot Filling

    2017

    Organized relational knowledge in the form of "knowledge graphs" is important for many applications. However, the ability to populate knowledge bases with facts automatically extracted from documents has improved frustratingly slowly. This paper simultaneously addresses …

  23. Simpler but More Accurate Semantic Dependency Parsing

    2018

    While syntactic dependency annotations concentrate on the surface or functional structure of a sentence, semantic dependency annotations aim to capture betweenword relationships that are more closely related to the meaning of a sentence, using graph-structured …

  24. Semi-Supervised Sequence Modeling with Cross-View Training

    2018

    Unsupervised representation learning algorithms such as word2vec and ELMo improve the accuracy of many supervised NLP models, mainly because they can take advantage of large amounts of unlabeled text. However, the supervised models only learn …

  25. Graph Convolution over Pruned Dependency Trees Improves Relation Extraction

    2018

    Dependency trees help relation extraction models capture long-range relations between words. However, existing dependency-based models either neglect crucial information (e.g., negation) by pruning the dependency trees too aggressively, or are computationally inefficient because it is …

  26. Universal Dependencies

    2021 · Computational Linguistics

    Abstract Universal dependencies (UD) is a framework for morphosyntactic annotation of human language, which to date has been used to create treebanks for more than 100 languages. In this article, we outline the linguistic theory …

  27. A Thorough Examination of the CNN/Daily Mail Reading Comprehension Task

    2016

    Enabling a computer to understand a document so that it can answer comprehension questions is a central, yet unsolved goal of NLP. A key factor impeding its solution by machine learned systems is the limited …

  28. Deep Biaffine Attention for Neural Dependency Parsing

    2016 · arXiv (Cornell University)

    In this paper, we present our approach for Multilingual Open Information Extraction. Our sequence labeling based approach builds only on Universal Dependency representation to capture OpenIE’s regularities and to perform Cross-lingual Multilingual OpenIE. We propose …

  29. Answering Complex Open-domain Questions Through Iterative Query Generation

    2019

    Peng Qi, Xiaowen Lin, Leo Mehr, Zijian Wang, Christopher D. Manning. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.

  30. ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators

    2020 · arXiv (Cornell University)

    Masked language modeling (MLM) pre-training methods such as BERT corrupt the input by replacing some tokens with [MASK] and then train a model to reconstruct the original tokens. While they produce good results when transferred …

  31. Emergent linguistic structure in artificial neural networks trained by self-supervision

    2020 · Proceedings of the National Academy of Sciences

    This paper explores the knowledge of linguistic structure learned by large artificial neural networks, trained via self-supervision, whereby the model simply tries to predict a masked word in a given context. Human language communication is …

  32. Stanza: A Python Natural Language Processing Toolkit for Many Human Languages

    2020

    We introduce Sta n z a , an open-source Python natural language processing toolkit supporting 66 human languages. Compared to existing widely used toolkits, Sta n z a features a language-agnostic fully neural pipeline for …