article Open access

Current and future state of evaluation of large language models for medical summarization tasks

  • npj Health Systems
Research footprint

At a glance

Citations
55
References
72
Comments
0
Paper overview

Abstract

Large Language Models have expanded the potential for clinical Natural Language Generation (NLG), presenting new opportunities to manage the vast amounts of medical text. However, their use in such high-stakes environments necessitate robust evaluation workflows. In this review, we investigated the current landscape of evaluation metrics for NLG in healthcare and proposed a future direction to address the resource constraints of expert human evaluation while balancing alignment with human judgments.

Record transparency

Publication details

DOI
10.1038/s44401-024-00011-2
OpenAlex
W4407108350
Document type
article
Language
EN
Source
npj Health Systems
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.