Arman Cohan
13 papers in the PaperMetrix corpus
Papers by this author
-
SciBERT: Pretrained Contextualized Embeddings for Scientific Text
2019
Obtaining large-scale annotated data for NLP tasks in the scientific domain is challenging and expensive. We release SciBERT, a pretrained language model based on BERT (Devlin et al., 2018) to address the lack of high-quality, …
-
SLEDGE: A Simple Yet Effective Baseline for Coronavirus Scientific Knowledge Search
2020 · arXiv (Cornell University)
With worldwide concerns surrounding the Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2), there is a rapidly growing body of literature on the virus. Clinicians, researchers, and policy-makers need a way to effectively search these articles. …
-
On Generating Extended Summaries of Long Documents
2020 · arXiv (Cornell University)
Prior work in document summarization has mainly focused on generating short summaries of a document. While this type of summary helps get a high-level view of a given document, it is desirable in some cases …
-
Cross-Document Language Modeling.
2021 · arXiv (Cornell University)
We introduce a new pretraining approach for language models that are geared to support multi-document NLP tasks. Our cross-document language model (CD-LM) improves masked language modeling for these tasks with two key ideas. First, we …
-
PRIMER: Pyramid-based Masked Sentence Pre-training for Multi-document Summarization.
2021 · arXiv (Cornell University)
Recently proposed pre-trained generation models achieve strong performance on single-document summarization benchmarks. However, most of them are pre-trained with general-purpose objectives and mainly aim to process single document inputs. In this paper, we propose PRIMER, …
-
Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization
2023 · arXiv (Cornell University)
While large language models (LLMs) can already achieve strong performance on standard generic summarization benchmarks, their performance on more complex summarization task settings is less studied. Therefore, we benchmark LLMs on instruction controllable text summarization, …
-
SciRIFF: A Resource to Enhance Language Model Instruction-Following over Scientific Literature
2024 · arXiv (Cornell University)
We present SciRIFF (Scientific Resource for Instruction-Following and Finetuning), a dataset of 137K instruction-following instances for training and evaluation, covering 54 tasks. These tasks span five core scientific literature understanding capabilities: information extraction, summarization, question …
-
YaleNLP @ PerAnsSumm 2025: Multi-Perspective Integration via Mixture-of-Agents for Enhanced Healthcare QA Summarization
2025
Automated summarization of healthcare community question-answering forums is challenging due to diverse perspectives presented across multiple user responses to each question.The PerAnsSumm Shared Task was therefore proposed to tackle this challenge by identifying perspectives from …
-
MSRS: Evaluating Multi-Source Retrieval-Augmented Generation
2025 · arXiv (Cornell University)
Retrieval-augmented systems are typically evaluated in settings where information required to answer the query can be found within a single source or the answer is short-form or factoid-based. However, many real-world applications demand the ability …
-
A Discourse-Aware Attention Model for Abstractive Summarization of Long Documents
2018
Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, Walter Chang, Nazli Goharian. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume …
-
CEDR
2019
Although considerable attention has been given to neural ranking architectures recently, far less attention has been paid to the term representations that are used as input to these models. In this work, we investigate how …
-
SciBERT: A Pretrained Language Model for Scientific Text
2019
Iz Beltagy, Kyle Lo, Arman Cohan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
-
Longformer: The Long-Document Transformer
2020 · arXiv (Cornell University)
The quadratic complexity of standard attention (O(N²)) remains the dominant bottleneck for training and deploying large language models on long sequences. We introduce Murmurative Attention, a novel attention mechanism that replaces pairwise token-token interactions with …