conference-paper
Open access
emrQA: A Large Corpus for Question Answering on Electronic Medical Records
Research footprint
At a glance
- Citations
- 179
- References
- 58
- Comments
- 0
Paper overview
Abstract
We propose a novel methodology to generate domain-specific large-scale question answering (QA) datasets by re-purposing existing annotations for other NLP tasks. We demonstrate an instance of this methodology in generating a large-scale QA dataset for electronic medical records by leveraging existing expert annotations on clinical notes for various NLP tasks from the community shared i2b2 datasets . The resulting corpus (emrQA) has 1 million questions-logical form and 400,000+ question-answer evidence pairs. We characterize the dataset and explore its learning potential by training baseline models for question to logical form and question to answer mapping.
Record transparency
Publication details
- DOI
- 10.18653/v1/d18-1258
- OpenAlex
- W2891113091
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.