A Multimodal Video Recommendation Method based on User-Embedded Graph Convolutional Networks and Fused Attention
At a glance
- الاستشهادات
- 0
- المراجع
- 8
- Comments
- 0
Abstract
The current Multi-modal Graph Convolutional Network(GCN) recommendation method only focuses on capturing the association and interaction between different modes in the representation learning of multi-modal item features. However, it embeds a large amount of preference-independent multi-modal noise, which contaminates the modeling of multi-modal user preferences. This approach also fails to consider the impact of user identity(ID) embedding on global graph structure information and modal information. Additionally, most existing multi-modal fusion methods only take into account the differences between modes, neglecting the potential sharing between them. In response to these limitations, this paper proposes a Multi-modal Video Recommendation method based on User-Embedded GCN and Fused Attention (MVR-UEGCN-FA). The method first purifies modal features using behavioral information from the project. It then embeds user ID to capture cooperation signals and semantic preferences separately before extending them to higher-order feature dimensions through graph convolution operations. Simultaneously, an improved attention mechanism is employed to extract common preferences across different modes and generate final recommendations. Experimental verification demonstrates that compared with baseline methods, the proposed method significantly improves recommendation performance in terms of recall rate index by 3.79%.
Publication details
- DOI
- 10.1145/3703935.3703993
- OpenAlex
- W4411730137
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.