conference-paper

Optimizing Multilingual Summarization: Advanced Corpus Creation and Filtering Approach

Research footprint

At a glance

Citations
1
References
24
Comments
0
Paper overview

Abstract

Summarization is essential for effectively understanding vast information, and machine learning is fundamental. to progressing this field. However, a major obstacle is the lack of high-quality datasets, especially for languages that are less commonly studied. To address this, we developed a large dataset on five low-resource languages in terms of input text and corresponding summary quality such as Arabic, Afrikaans, Azerbaijani, Albanian, and Bengali. Here, we commenced by translating texts and summaries from the prevalent Daily Mail CNN dataset into these languages. To ensure the quality and relevance, we introduced a filtering approach by using robust metrics such as LaBSE, ROUGE, BLEU, and BERTScore to filter the data, eliminating outliers and ensuring quality. This procedure yielded a substantial dataset comprising 567,462 varied and high-quality instances, with each language containing more than 100,000 text-summary pairs.Furthermore, this study uses extensive comparative analysis, focusing on all five languages to demonstrate that the abstractive summarization model performs better compared to the existing work. The trained MT5-base model on our developed dataset outperforms the state-of-the-art by 7% in BERTScores, 8% in BLEU, and 16% in ROUGE.While highlighting the importance of high-quality datasets and models, this work advances abstractive multilingual summarization.

Record transparency

Publication details

DOI
10.1109/iccit64611.2024.11022549
OpenAlex
W4411172320
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.