article Open access

Improving text classification via computing category correlation matrix from text graph

  • Computer Speech & Language
  • Elsevier BV
Research footprint

At a glance

Citations
3
References
39
Comments
0
Paper overview

Abstract

In text classification task , models have shown remarkable accuracy across various datasets. However, confusion often arises when certain categories within the dataset are too similar, causing misclassification of certain samples. This paper proposes an improved method for this problem, through the creation of a three-layer text graph for the corpus, which is used to calculate the Category Correlation Matrix (CCM). Additionally, this paper introduces category-adaptive contrastive learning for text embedding from the encoder, enhancing the model’s ability to distinguish between samples in confusable categories that are easily confused. Soft labels are generated using this matrix to guide the classifier, preventing the model from becoming overconfident with one-hot vectors. The efficacy of this approach was demonstrated through experimental evaluations on three text encoders and six different datasets.

Record transparency

Publication details

DOI
10.1016/j.csl.2024.101688
OpenAlex
W4400453544
Document type
article
Language
EN
Source
Computer Speech & Language
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.