Revisiting Spatial Inductive Bias with MLP-Like Model
At a glance
- الاستشهادات
- 0
- المراجع
- 23
- Comments
- 0
Abstract
In recent years, deep convolution neural nets have produced outstanding results in vision tasks compared to previous methods. In the Transformer models with self-attention structure, the introduction of inductive bias of convolution operator has been studied. There is no doubt that the hard inductive bias of the convolution operator, i.e., the receptive field (locality) and spatial invariance (weight sharing) in the spatial dimension, is an essential factor in achieving high performance. We hypothesize that by placing additional constraints on the locality, we can obtain better inductive bias. In this paper, we propose an MLP-like model, content-aware token mixing MLP (CaMLP), in which locality and spatial invariance are tunable. We verify that smooth and outward-decaying locality plays an important role in token mixing of the model with relaxed spatial invariance in the CIFAR-10/ImageNet100 classification task.
Publication details
- DOI
- 10.1109/icip46576.2022.9897403
- OpenAlex
- W4308237010
- Document type
- conference-paper
- Language
- EN
- Source
- 2022 IEEE International Conference on Image Processing (ICIP)
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.