article

Flexpéro: Flexible Expressive Zero-Shot Speech Refinement via In-Context Learning

  • IEEE Signal Processing Letters
  • Institute of Electrical and Electronics Engineers
Research footprint

At a glance

Citations
0
References
37
Comments
0
Paper overview

Abstract

Controlling speech expressiveness has emerged as a critical research frontier in speech generation, focusing on synthesizing natural, human-like speech that accurately conveys intended psychological and emotional states. While many large-scale models have demonstrated sufficient zero-shot capability by conditioning the acoustic model on reference speech—which provides cues on speaker identity and style—they often fall short of meeting desired emotional or prosodic targets at fine-grained levels. To address this challenge, we propose a novel speech refinement method based on a zero-shot voice synthesis model that can flexibly and interactively enhance expressiveness on unsatisfactory speech segments. It supports emotion modulation through chunk-wise valence/arousal and flexible keyframe-based prediction of pitch and energy, allowing for the creation of any prosodic patterns. Experimental results show that our method achieves fine-grained control, thus enriching the expressiveness of zero-shot synthetic speech.

Record transparency

Publication details

DOI
10.1109/lsp.2025.3586183
OpenAlex
W4412081436
Document type
article
Language
EN
Source
IEEE Signal Processing Letters
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.