conference-paper

GSWIN: Gated MLP Vision Model with Hierarchical Structure of Shifted Window

Research footprint

At a glance

Citations
6
References
44
Comments
0
Paper overview

Abstract

Following the success in language domain, the self-attention mechanism (Transformer) has been adopted in the vision domain and achieving great success recently. Additionally, the use of multi-layer perceptron (MLP) is also explored in the vision domain as another stream. These architectures have been attracting attention recently to alternate the traditional CNNs, and many Vision Transformers and Vision MLPs have been proposed. By fusing the above two streams, this paper proposes gSwin, a novel vision model which can consider spatial hierarchy and locality due to its network structure similar to the Swin Transformer, and is parameter efficient due to its gated MLP-based architecture. It is experimentally confirmed that the gSwin can achieve better accuracy than Swin Transformer on three common tasks of vision with smaller model size.

Record transparency

Publication details

DOI
10.1109/icassp49357.2023.10096453
OpenAlex
W4372267514
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.