preprint Open access

A Chinese Short Text Similarity Method Integrating Sentence-level and Phrase-level Semantics

  • Preprints.org
Research footprint

At a glance

Citations
1
References
0
Comments
0
Paper overview

Abstract

Short text similarity, as a pivotal research domain within Natural Language Processing (NLP), has been extensively utilized in intelligent search, recommendation systems, and question-answering systems. The majority of existing models for short text similarity concentrate on aligning the overall semantic content of entire sentences, frequently neglecting the semantic correlations between individual phrases within the sentences. This challenge is particularly acute in the Chinese language context, where synonyms and near-synonyms can introduce substantial interference in the computation of text similarity. In this paper, we introduce a short text similarity computation methodology that integrates both sentence-level and phrase-level semantics. By harnessing vector representations of Chinese words/phrases as external knowledge, our approach amalgamates global sentence characteristics with local phrase features to compute short text similarity from diverse perspectives, spanning from the global to the local level. Experimental findings substantiate that the proposed model surpasses previous approaches in Chinese short text similarity tasks. Specifically, it attains an accuracy of 90.16% on the LCQMC, marking an enhancement of 2.23% over ERNIE and 1.46% over the previously top-performing model, Glyce + BERT.

Record transparency

Publication details

DOI
10.20944/preprints202411.0453.v1
OpenAlex
W4404216636
Document type
preprint
Language
EN
Source
Preprints.org
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.