Intermediate-Task Learning with Pretrained Model for Synthesized Speech MOS Prediction
At a glance
- الاستشهادات
- 1
- المراجع
- 30
- Comments
- 0
Abstract
Mean Opinion Scores (MOS) prediction has been a crucial task for speech quality assessment. Recently, methods based on pre-training and fine-tuning achieved the start-of-the-art performance for MOS prediction. While these two-stage methods can improve the generalization ability of the model, there still exists a large gap between the pre-training and the downstream MOS prediction objectives in the input distribution and the output label space. In this paper, we proposed a three-stage intermediate task training scheme, then we tailored two possible intermediate tasks (classification and contrastive learning) for the speech MOS prediction task. The experimental results show that the proposed models achieve the SOTA results compared with the existing two-stage systems in most of the metrics for in-domain and out-of-domain datasets with less computational complexity.
Publication details
- DOI
- 10.1109/icme55011.2023.00072
- OpenAlex
- W4386159905
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.