review

Generative Large Language Models for Question Answering from Electronic Health Records: A Systematic Review

  • Open MIND
Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Abstract

This systematic review synthesises evidence on the performance, methodological quality, and clinical applicability of generative large language models for question-answering tasks using electronic health record data. We searched PubMed, Scopus, Web of Science, IEEE Xplore, and ACM Digital Library for original empirical studies evaluating decoder-only or encoder-decoder transformer architectures applied to clinical question-answering over EHR data. Risk of bias was assessed using an adapted QUADAS-2 framework with signalling questions modified to address LLM-specific methodological concerns including data contamination, LLM-as-judge bias, and reproducibility. Due to anticipated heterogeneity in task definitions and evaluation metrics, findings are synthesised narratively rather than through meta-analysis. A preliminary clinical implementation decision framework is developed to translate research findings into actionable guidance for healthcare institutions evaluating these systems for deployment.

Record transparency

Publication details

OpenAlex
W7154672257
Document type
review
Language
EN
Source
Open MIND
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.