Multi-Class Token Attention Learning for Weakly Supervised Semantic Segmentation
At a glance
- الاستشهادات
- 0
- المراجع
- 19
- Comments
- 0
Abstract
Weakly supervised semantic segmentation (WSSS) using only image-level class labels is a challenging task. Due to the local receptive fields of convolutional neural networks (CNNs), CAMs applied to CNNs often suffer from partial activation-highlighting the most discriminative parts of the image. To capture both local features and global representations, we propose an approach called MCTRM, which is a WSSS solution based on an improved Conformer as backbone. To ensure each class token's discriminative ability for the respective target class, we design a class-aware training strategy (CAT) to establish a one-to-one correspondence between the output class tokens and the ground-truth class labels. To enhance the class discrimination ability between each class token, we propose a class token regularization module (CTRM) to enhance the learning of discriminative class tokens so that the model can better capture the unique features and attributes of each class. To obtain a fine-grained CAM, we propose the refinement module (RM) to combine the class-specific attention and inherent patch-level pairwise affinity between image patches learned by the Transformer branch with the CAM of the CNN branch. Despite its simplicity, MCTRM achieves the state-of-the-art performance of 69.0% and 68.4% on the PASCAL VOC 2012 validation and test sets, respectively.
Publication details
- DOI
- 10.1109/icmlc63072.2024.10935262
- OpenAlex
- W4408897238
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.