Beyond Label: Cold-Start Self-Training for Cross-Domain Semantic Text Similarity
At a glance
- Citations
- 0
- References
- 24
- Comments
- 0
Öz
In Natural Language Processing (NLP), comprehending the semantic connection between two texts, a Semantic Text Similarity (STS) task, poses a significant challenge. This challenge is especially pronounced in resource-constrained and cross-domain contexts, where traditional methods are hindered by the high costs associated with data labeling. We propose an innovative technique, designated as “Cold-Start Self-Training” that reduces reliance on large labeled datasets for STS tasks in resource-restricted settings. This method utilizes dual-view pooling to extract semantic similarity information from unlabeled data and generates pseudo-labeled data to fine-tune the cross-encoder model. Dual-view pooling combines different pooling results of the same text to evaluate semantic similarity without additional model tuning, simplifying the self-training process. Experimental results show that our method significantly improves the cross-encoder model's performance on STS tasks in the medical domain. Our findings provide new strategies for cross-domain STS tasks, challenging the traditional reliance on extensive labeled data. We also validate the potential of unsupervised pretrained models for cross-domain tasks, offering theoretical and practical support for complex challenges.
Publication details
- DOI
- 10.1109/smc54092.2024.10831873
- OpenAlex
- W4406612301
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.