Comparative Evaluation of Language Models in Summarizing Chinese Healthcare Exam Findings
At a glance
- Citations
- 0
- References
- 33
- Comments
- 0
Abstract
Abstract Background and Objectives: Summarizing Chinese healthcare exam findings (SCHEF) is a critical and challenging task that requires extensive medical knowledge. Despite its importance, research on automating the generation of such summaries remains limited.We aim to develop a standardized Chinese Healthcare Exam Summarization (CHES) dataset and evaluate the performance of various sequence-to-sequence(seq2seq) models on the task of summarizing overall conclusions of healthcare exam findings. Our primary focus is on assessing the models’applicability, generalization, and transferability in this context. Materials and Methods: We collect healthcare exam data from over 110, 000 individuals across two centers, creating the anonymized and cleaned Chinese Healthcare Exam Summarization (CHES) dataset for training and evaluating summarization models. We conduct comprehensive experiments on a broad range of summarization methods, including the manual extraction baseline method, a variety of RNN models (RNN, GRU, LSTM), PointerNet with a copy mechanism, convolutional seq2seq techniques, and self-attention-driven architectures such as Transformer, RoBERTa, and MacBERT. Additionally, we explore the significance of domain adaptation for BERT-based models on specific task datasets. Results: Our experiments show that BERT-based models demonstrate the best performance in generating summaries for Chinese healthcare exam findings, as evidenced by F1 and recall evaluation metrics. Through this domain adaptation on task-specific datasets, we notice a substantial improvement of approximately 3% in the generated text summarization. Conclusion: Our study demonstrates that advanced seq2seq language models can be readily employed for generating Chinese healthcare exam findings summarizations without requiring substantial modifications. Further pretraining language models on target datasets considerably improves summary quality by enhancing adherence to medical semantic standards, as it enables the model to better capture contextual and semantic information.
Publication details
- DOI
- 10.21203/rs.3.rs-3847983/v1
- OpenAlex
- W4390765576
- Document type
- preprint
- Language
- EN
- Source
- Research Square
- Last metadata update
Comments
Log in to join the discussion.