conference-paper

Multi-Class Token Attention Learning for Weakly Supervised Semantic Segmentation

Research footprint

At a glance

Citations
0
References
19
Comments
0
Paper overview

Öz

Weakly supervised semantic segmentation (WSSS) using only image-level class labels is a challenging task. Due to the local receptive fields of convolutional neural networks (CNNs), CAMs applied to CNNs often suffer from partial activation-highlighting the most discriminative parts of the image. To capture both local features and global representations, we propose an approach called MCTRM, which is a WSSS solution based on an improved Conformer as backbone. To ensure each class token's discriminative ability for the respective target class, we design a class-aware training strategy (CAT) to establish a one-to-one correspondence between the output class tokens and the ground-truth class labels. To enhance the class discrimination ability between each class token, we propose a class token regularization module (CTRM) to enhance the learning of discriminative class tokens so that the model can better capture the unique features and attributes of each class. To obtain a fine-grained CAM, we propose the refinement module (RM) to combine the class-specific attention and inherent patch-level pairwise affinity between image patches learned by the Transformer branch with the CAM of the CNN branch. Despite its simplicity, MCTRM achieves the state-of-the-art performance of 69.0% and 68.4% on the PASCAL VOC 2012 validation and test sets, respectively.

Record transparency

Publication details

DOI
10.1109/icmlc63072.2024.10935262
OpenAlex
W4408897238
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.