article وصول مفتوح

Disentangling Semantic-to-Visual Confusion for Zero-Shot Learning

  • IEEE Transactions on Multimedia
  • Institute of Electrical and Electronics Engineers
Research footprint

At a glance

الاستشهادات
26
المراجع
57
Comments
0
Paper overview

Abstract

Using generative models to synthesize visual features from semantic distribution is one of the most popular solutions to ZSL image classification in recent years. The triplet loss (TL) is popularly used to generate realistic visual distributions from semantics by automatically searching discriminative representations. However, the traditional TL cannot search reliable unseen disentangled representations due to the unavailability of unseen classes in ZSL. To alleviate this drawback, we propose in this work a multi-modal triplet loss (MMTL) which utilizes multi-modal information to search adisentangledrepresentation space. As such, all classes can interplay which can benefit learning disentangled class representations in the searched space. Furthermore, we develop a novel model called Disentangling Class Representation Generative Adversarial Network (DCR-GAN) focusing on exploiting the disentangled representations in training, feature synthesis, and final recognition stages. Benefiting from the disentangled representations, DCR-GAN could fit a more realistic distribution over both seen and unseen features. Extensive experiments show that our proposed model can lead to superior performance to the state-of-the-arts on four benchmark datasets.

Record transparency

Publication details

DOI
10.1109/tmm.2021.3089017
OpenAlex
W3169149706
Document type
article
Language
EN
Source
IEEE Transactions on Multimedia
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.