Research on Text Generated Images Based on Deep Fusion Attention
At a glance
- Citations
- 0
- References
- 7
- Comments
- 0
Öz
This article proposes a deep fusion attention generative adversarial network (DFAttn-GAN) to address the issues of insufficient text image semantic alignment and weak ability to generate complex scene details in text to image generation tasks. This model has undergone key optimizations based on the AttnGAN framework to improve generation quality: (1) introducing a deep fusion attention mechanism, adding a global consistency loss in the third stage, generating images through contrastive learning constraints to match input text more accurately, while reducing similarity with unmatched samples; (2)By introducing a single discriminator to determine the overall quality of the generated image, the training$p$rocess has been significantly improved in terms of stability and convergence speed. The performance was validated on the CUB-200-2011 and MS-COCO datasets, and the results showed that DFAttn-GAN outperformed the DF-GAN baseline model in both Inception Score and FID metrics. This study provides an effective solution for high-resolution and high semantic consistency text to image generation tasks, with broad potential applications in fields such as virtual scene construction and art design.
Publication details
- DOI
- 10.1109/prmvai65741.2025.11108650
- OpenAlex
- W4413278202
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.