Daniel S. Weld
7 papers in the PaperMetrix corpus
Papers by this author
-
High-Precision Extraction of Emerging Concepts from Scientific Literature
2020
Identification of new concepts in scientific literature can help power faceted search, scientific trend analysis, knowledge-base construction, and more, but current methods are lacking. Manual identification can't keep up with the torrent of new publications, …
-
Qlarify: Recursively Expandable Abstracts for Dynamic Information Retrieval over Scientific Papers
2024
Navigating the vast scientific literature often starts with browsing a paper’s abstract. However, when a reader seeks additional information, not present in the abstract, they face a costly cognitive chasm during their dive into the …
-
Fine-Grained Entity Recognition
2021 · Proceedings of the AAAI Conference on Artificial Intelligence
Entity Recognition (ER) is a key component of relation extraction systems and many other natural-language processing applications. Unfortunately, most ER systems are restricted to produce labels from to a small set of entity classes, e.g., …
-
TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
2017
We present TriviaQA, a challenging reading comprehension dataset containing over 650K question-answer-evidence triples. TriviaQA includes 95K questionanswer pairs authored by trivia enthusiasts and independently gathered evidence documents, six per question on average, that provide high …
-
StaQC
2018
Stack Overflow (SO) has been a great source of natural language questions and their code solutions (i.e., question-code pairs), which are critical for many tasks including code retrieval and annotation. In most existing research, question-code …
-
SpanBERT: Improving Pre-training by Representing and Predicting Spans
2020 · Transactions of the Association for Computational Linguistics
We present SpanBERT, a pre-training method that is designed to better represent and predict spans of text. Our approach extends BERT by (1) masking contiguous random spans, rather than random tokens, and (2) training the …
-
S2ORC: The Semantic Scholar Open Research Corpus
2020
We introduce S2ORC, 1 a large corpus of 81.1M English-language academic papers spanning many academic disciplines. The corpus consists of rich metadata, paper abstracts, resolved bibliographic references, as well as structured full text for 8.1M …