Self-Supervised Text-Centric Sentiment Analysis via Sparse Cross-Modal Interactions
At a glance
- Citations
- 0
- References
- 14
- Comments
- 0
Abstract
Multi-modal sentiment analysis serves as a crucial research sub-topic within the field of multi-modal learning. Significant advancements have been achieved with the emergence of transformers and attention mechanisms, yet several prominent issues remain. Firstly, existing methods overlook the imbalanced contributions among modalities, treating the three modalities (typically text, audio, and visual) as equals without acknowledging the paramount importance of the text modality. Secondly, these methods neglect the redundancy and noise generated during cross-modal interactions. Additionally, they tend to focus on learning modal consistency information while disregarding the essential learning of modal disparity information. To address these issues, this paper proposes Self-Supervised Text-Centric Sentiment Analysis via Sparse Cross-Modal Interactions (SST-SMI). This text-centric architecture leverages text to guide the learning of less resourceful non-text modalities, fully utilizing the highest contribution of the text modality. A Top-K sparse attention transformer is employed as a means of modal interaction, suppressing inter-modal noise during the interaction process. Furthermore, the introduction of the Unimodal Label Generation Module preserves the disparity information of each modality. Extensive experiments conducted on the CMU-MOSI and CMU-MOSEI datasets demonstrate that SST-SMI outperforms existing methods.
Publication details
- DOI
- 10.1109/iccbd-ai65562.2024.00021
- OpenAlex
- W4408860450
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.