Researcher profile

Daniel S. Weld

7 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. High-Precision Extraction of Emerging Concepts from Scientific Literature

    2020

    Identification of new concepts in scientific literature can help power faceted search, scientific trend analysis, knowledge-base construction, and more, but current methods are lacking. Manual identification can't keep up with the torrent of new publications, …

  2. Qlarify: Recursively Expandable Abstracts for Dynamic Information Retrieval over Scientific Papers

    2024

    Navigating the vast scientific literature often starts with browsing a paper’s abstract. However, when a reader seeks additional information, not present in the abstract, they face a costly cognitive chasm during their dive into the …

  3. Fine-Grained Entity Recognition

    2021 · Proceedings of the AAAI Conference on Artificial Intelligence

    Entity Recognition (ER) is a key component of relation extraction systems and many other natural-language processing applications. Unfortunately, most ER systems are restricted to produce labels from to a small set of entity classes, e.g., …

  4. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

    2017

    We present TriviaQA, a challenging reading comprehension dataset containing over 650K question-answer-evidence triples. TriviaQA includes 95K questionanswer pairs authored by trivia enthusiasts and independently gathered evidence documents, six per question on average, that provide high …

  5. StaQC

    2018

    Stack Overflow (SO) has been a great source of natural language questions and their code solutions (i.e., question-code pairs), which are critical for many tasks including code retrieval and annotation. In most existing research, question-code …

  6. SpanBERT: Improving Pre-training by Representing and Predicting Spans

    2020 · Transactions of the Association for Computational Linguistics

    We present SpanBERT, a pre-training method that is designed to better represent and predict spans of text. Our approach extends BERT by (1) masking contiguous random spans, rather than random tokens, and (2) training the …

  7. S2ORC: The Semantic Scholar Open Research Corpus

    2020

    We introduce S2ORC, 1 a large corpus of 81.1M English-language academic papers spanning many academic disciplines. The corpus consists of rich metadata, paper abstracts, resolved bibliographic references, as well as structured full text for 8.1M …