article Open access

Lexicon‐based fine‐tuning of multilingual language models for low‐resource language sentiment analysis

  • CAAI Transactions on Intelligence Technology
  • Institution of Engineering and Technology
Research footprint

At a glance

Citations
14
References
32
Comments
0
Paper overview

Abstract

Abstract Pre‐trained multilingual language models (PMLMs) such as mBERT and XLM‐R have shown good cross‐lingual transferability. However, they are not specifically trained to capture cross‐lingual signals concerning sentiment words. This poses a disadvantage for low‐resource languages (LRLs) that are under‐represented in these models. To better fine‐tune these models for sentiment classification in LRLs, a novel intermediate task fine‐tuning (ITFT) technique based on a sentiment lexicon of a high‐resource language (HRL) is introduced. The authors experiment with LRLs Sinhala, Tamil and Bengali for a 3‐class sentiment classification task and show that this method outperforms vanilla fine‐tuning of the PMLM. It also outperforms or is on‐par with basic ITFT that relies on an HRL sentiment classification dataset.

Record transparency

Publication details

DOI
10.1049/cit2.12333
OpenAlex
W4393374551
Document type
article
Language
EN
Source
CAAI Transactions on Intelligence Technology
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.