FGCSA-Net: A Novel Framework for Medical Report Generation Via Fine-Grained Feature Preservation and Semantic Alignment
At a glance
- الاستشهادات
- 0
- المراجع
- 0
- Comments
- 0
Abstract
Medical image report generation is a cross-modal task that converts medical images into structured radiology reports. Existing methods often struggle with two issues: loss of subtle visual details in deep networks and weak alignment between image features and report semantics. To address these problems, we propose Fine-Grained Cross-modal Semantic Alignment Network (FGCSA-Net), which combines residual feature preservation with cross-attention-based visual-text alignment in a large-language-model framework. Residual connections preserve diagnostically important local details during visual encoding, while the cross-attention module links textual decoding to the most relevant image regions. Experiments on MIMIC-CXR and IU-Xray show that FGCSA-Net improves report generation quality over strong baselines; in particular, ROUGE-L improves by 26.7% compared with XrayGPT.
Publication details
- DOI
- 10.1109/jbhi.2026.3710597
- OpenAlex
- W7167464092
- Document type
- article
- Language
- EN
- Source
- IEEE Journal of Biomedical and Health Informatics
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.