conference-paper

Multi-modal sentiment analysis based on expression and text

Research footprint

At a glance

الاستشهادات
0
المراجع
11
Comments
0
Paper overview

Abstract

Sentiment analysis has always been a hot research topic in the field of natural language processing. With the development of artificial intelligence and big data technology, many data such as text, images, and videos that have personal emotional tendencies are accidentally generated. Initially, researchers mainly focused on studying text data, but over time, more and more people began to realize the limitations of single-mode analysis. By introducing multi-modal data such as images and videos, more dimensional information can be provided for sentiment analysis, thereby improving performance. Therefore, more and more researchers are exploring richer sources of emotional information to improve the effectiveness of sentiment analysis. This article proposes a multi-modal sentiment analysis method based on deep learning and designs an experimental model targeted towards video modal data. In terms of text, this article proposes a GloVe-based BiGRU word vector model to process data. For images, the method changes from using simple frame extraction to speaker anchoring, uses object detection technology to segment the region of interest for the subsequent model to extract image expression features. A pre-trained ResNet101 model is used to obtain vectors and generate a sequenced image matrix, which is then input to LSTM for processing. From the experimental results, it can be seen that this method performs better than existing methods in terms of accuracy and F1 scores.

Record transparency

Publication details

DOI
10.1117/12.3004687
OpenAlex
W4387452724
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.