article

Flexpéro: Flexible Expressive Zero-Shot Speech Refinement via In-Context Learning

  • IEEE Signal Processing Letters
  • Institute of Electrical and Electronics Engineers
Research footprint

At a glance

الاستشهادات
0
المراجع
37
Comments
0
Paper overview

Abstract

Controlling speech expressiveness has emerged as a critical research frontier in speech generation, focusing on synthesizing natural, human-like speech that accurately conveys intended psychological and emotional states. While many large-scale models have demonstrated sufficient zero-shot capability by conditioning the acoustic model on reference speech—which provides cues on speaker identity and style—they often fall short of meeting desired emotional or prosodic targets at fine-grained levels. To address this challenge, we propose a novel speech refinement method based on a zero-shot voice synthesis model that can flexibly and interactively enhance expressiveness on unsatisfactory speech segments. It supports emotion modulation through chunk-wise valence/arousal and flexible keyframe-based prediction of pitch and energy, allowing for the creation of any prosodic patterns. Experimental results show that our method achieves fine-grained control, thus enriching the expressiveness of zero-shot synthetic speech.

Record transparency

Publication details

DOI
10.1109/lsp.2025.3586183
OpenAlex
W4412081436
Document type
article
Language
EN
Source
IEEE Signal Processing Letters
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.