conference-paper

Multimodal Taste Classification of Chinese Recipe Based on Image and Text Fusion

  • 2020 5th International Conference on Smart Grid and Electrical Automation (ICSGEA)
Research footprint

At a glance

الاستشهادات
4
المراجع
17
Comments
0
Paper overview

Abstract

It is difficult for taste classification of Chinese recipe to achieve satisfactory results based on single-modal data. However, there are few studies on multimodal analysis in this field. In this paper, we put forward to tackle taste classification for Chinese recipe based on image and text fusion algorithms. Firstly, visual features and textual features are extracted from different models, including a convolutional neural network (CNN) constructed for visual feature extraction and a pretrained word2vec model combined with a multi-layer perception network for textual feature extraction. Secondly, two fusion strategies, called feature-level and decision-level fusion, are designed to perform multimodal fusion for the final taste prediction. Several experiments are carried out with K-fold cross-validation to verify the effectiveness of our proposed model. The results show that the multimodal fusion model for taste classification is superior to those based on single-modal features. Besides, compared with feature-level fusion, decision-level fusion performs better in the task of taste classification for Chinese recipe.

Record transparency

Publication details

DOI
10.1109/icsgea51094.2020.00050
OpenAlex
W3107706539
Document type
conference-paper
Language
EN
Source
2020 5th International Conference on Smart Grid and Electrical Automation (ICSGEA)
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.