A Fused Transformer-Based Ensemble Approach Using Pre-Trained LLMs for Multi-Modal News Summarization
At a glance
- Citations
- 1
- References
- 21
- Comments
- 0
Abstract
News agencies and media outlets brace for impact as information overload from several sectors risks undermining confidence in sectorial readership. News summarization filters information and ensures that readers have relevant news updates. Combining texts, images, and videos for summarizing news is an efficient approach. However, there exist challenges while integrating multi-modal information in learning models - i.e., learning models must find relevant information, successfully fuse features, and close the semantic gap between these multi-modalities. The multi-modal news summarization aspect can lead to biases or low accuracy while processing large volumes of dynamic news content in real time. In this paper, a fused ensemble approach combining multiple Pre-trained Language Models (PLMs) is designed for the multimodal news summarization problem. It involves Large Language Models (LLMs) supported with Bootstrapped Language-Image Pretraining (BLIP-2) and Bidirectional Encoder Representational from Transformers (BERT) to effectively deal with text and visual elements of news articles. In addition, a comparative study of the proposed fused transformer-based ensemble approach with a few baseline learning methods such as traditional LLMs is carried out at the IoT Cloud research laboratory. The results revealed a higher grammatical mean value of 0.77 and 0.74 quality in summarizing news contents from various input sources. The proposed approach will be useful for magazine readership or sector-specific readership when tailored to their specific problem domains.
Publication details
- DOI
- 10.1109/icict64420.2025.11004965
- OpenAlex
- W4410639820
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.