conference-paper

Research on Text Generated Images Based on Deep Fusion Attention

Research footprint

At a glance

Citations
0
References
7
Comments
0
Paper overview

Öz

This article proposes a deep fusion attention generative adversarial network (DFAttn-GAN) to address the issues of insufficient text image semantic alignment and weak ability to generate complex scene details in text to image generation tasks. This model has undergone key optimizations based on the AttnGAN framework to improve generation quality: (1) introducing a deep fusion attention mechanism, adding a global consistency loss in the third stage, generating images through contrastive learning constraints to match input text more accurately, while reducing similarity with unmatched samples; (2)By introducing a single discriminator to determine the overall quality of the generated image, the training$p$rocess has been significantly improved in terms of stability and convergence speed. The performance was validated on the CUB-200-2011 and MS-COCO datasets, and the results showed that DFAttn-GAN outperformed the DF-GAN baseline model in both Inception Score and FID metrics. This study provides an effective solution for high-resolution and high semantic consistency text to image generation tasks, with broad potential applications in fields such as virtual scene construction and art design.

Record transparency

Publication details

DOI
10.1109/prmvai65741.2025.11108650
OpenAlex
W4413278202
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.