conference-paper

Self-Supervised Text-Centric Sentiment Analysis via Sparse Cross-Modal Interactions

Research footprint

At a glance

الاستشهادات
0
المراجع
14
Comments
0
Paper overview

Abstract

Multi-modal sentiment analysis serves as a crucial research sub-topic within the field of multi-modal learning. Significant advancements have been achieved with the emergence of transformers and attention mechanisms, yet several prominent issues remain. Firstly, existing methods overlook the imbalanced contributions among modalities, treating the three modalities (typically text, audio, and visual) as equals without acknowledging the paramount importance of the text modality. Secondly, these methods neglect the redundancy and noise generated during cross-modal interactions. Additionally, they tend to focus on learning modal consistency information while disregarding the essential learning of modal disparity information. To address these issues, this paper proposes Self-Supervised Text-Centric Sentiment Analysis via Sparse Cross-Modal Interactions (SST-SMI). This text-centric architecture leverages text to guide the learning of less resourceful non-text modalities, fully utilizing the highest contribution of the text modality. A Top-K sparse attention transformer is employed as a means of modal interaction, suppressing inter-modal noise during the interaction process. Furthermore, the introduction of the Unimodal Label Generation Module preserves the disparity information of each modality. Extensive experiments conducted on the CMU-MOSI and CMU-MOSEI datasets demonstrate that SST-SMI outperforms existing methods.

Record transparency

Publication details

DOI
10.1109/iccbd-ai65562.2024.00021
OpenAlex
W4408860450
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.