article وصول مفتوح

Lightweight open-source large language models versus cTAKES for information extraction from discharge summaries: tobacco smoking status test case

  • JAMIA Open
  • University of Oxford
Research footprint

At a glance

الاستشهادات
0
المراجع
17
Comments
0
Paper overview

Abstract

Abstract Objectives To compare lightweight open-source large language models (LLMs) with cTAKES, a state-of-the-art natural language processing (NLP) system, in an information extraction task from hospitalization discharge summaries. Materials and Methods Two readers annotated 250 randomly sampled adult discharge summaries (BJC HealthCare, 2018-2023) for tobacco smoking status as “Smoker,” “Never smoker,” “Unknown.” Six LLMs (Llama-3 [1B-70B], gpt-oss-20B, MedGemma-27B) and cTAKES extracted smoking status from summaries. Performance was benchmarked against consensus annotations using weighted F1-score, macro F1-score, and per-class F1-scores and a noninferiority test. Results Inter-reader agreement was excellent (κ = 0.91). LLM size (2.3-47.3 GB) and inference time (2.5-14.5 s/note) varied. gpt-oss-20B achieved non-inferior performance vs cTAKES (F1 = 0.99 vs 0.97; P < .021). Discussion The high accuracy and efficiency of gpt-oss-20B support its potential as a practical, open-source alternative to traditional NLP for clinical information extraction. Conclusion Lightweight LLMs can be applied for use across diverse clinical information extraction tasks without the need for task-specific fine-tuning.

Record transparency

Publication details

DOI
10.1093/jamiaopen/ooaf182
OpenAlex
W7124245478
Document type
article
Language
EN
Source
JAMIA Open
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.