conference-paper Open access

Cross Modal Learning Method for Multimodal Discourse Generation and Understanding

  • Procedia Computer Science
  • Elsevier BV
Research footprint

At a glance

Citations
0
References
12
Comments
0
Paper overview

Abstract

To solve the problem that traditional single-modal methods cannot make full use of multimodal information, it is necessary to use multimodal methods. To this end, this paper adopted a cross modal learning method to comprehensively use the information of multiple modes in view of the problem of multimodal discourse generation and understanding. Through the modeling of the relationship between modals and the fusion of features, a cross modal learning model was constructed. The model was deeply trained using the maximum quasi-probability estimation, and experimental evaluation was carried out. The experimental results showed that the accuracy rate of this method was between 89% and 96%, and the cross modal learning method of multimodal discourse generation and understanding showed obvious advantages in the average accuracy rate. Compared with the traditional single-mode method, the cross modal learning method can provide more accurate and rich model prediction and generation capabilities.

Record transparency

Publication details

DOI
10.1016/j.procs.2026.06.003
OpenAlex
W7167703146
Document type
conference-paper
Language
EN
Source
Procedia Computer Science
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.