Flexpéro: Flexible Expressive Zero-Shot Speech Refinement via In-Context Learning
At a glance
- Citations
- 0
- References
- 37
- Comments
- 0
Abstract
Controlling speech expressiveness has emerged as a critical research frontier in speech generation, focusing on synthesizing natural, human-like speech that accurately conveys intended psychological and emotional states. While many large-scale models have demonstrated sufficient zero-shot capability by conditioning the acoustic model on reference speech—which provides cues on speaker identity and style—they often fall short of meeting desired emotional or prosodic targets at fine-grained levels. To address this challenge, we propose a novel speech refinement method based on a zero-shot voice synthesis model that can flexibly and interactively enhance expressiveness on unsatisfactory speech segments. It supports emotion modulation through chunk-wise valence/arousal and flexible keyframe-based prediction of pitch and energy, allowing for the creation of any prosodic patterns. Experimental results show that our method achieves fine-grained control, thus enriching the expressiveness of zero-shot synthetic speech.
Publication details
- DOI
- 10.1109/lsp.2025.3586183
- OpenAlex
- W4412081436
- Document type
- article
- Language
- EN
- Source
- IEEE Signal Processing Letters
- Last metadata update
Comments
Log in to join the discussion.