Researcher profile

Zhiyong Wu

14 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Applying Multitask Learning to Acoustic-Phonemic Model for Mispronunciation Detection and Diagnosis in L2 English Speech

    2018

    For mispronunciation detection and diagnosis (MDD), nowadays approaches generally treat the phonemes in correct and mispronunciations as the same despite the fact they may actually carry different characteristics. Furthermore, serious data imbalance issue between correct …

  2. Speech-XLNet: Unsupervised Acoustic Model Pretraining for Self-Attention Networks

    2020

    Self-attention network (SAN) can benefit significantly from the bi-directional representation learning through unsupervised pretraining paradigms such as BERT and XLNet.In this paper, we present an XLNet-like pretraining scheme "Speech-XLNet" to learn speech representations with self-attention …

  3. FERNet: Fine-grained Extraction and Reasoning Network for Emotion Recognition in Dialogues

    2020

    Yingmei Guo, Zhiyong Wu, Mingxing Xu. Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing. 2020.

  4. Controllable Emphatic Speech Synthesis based on Forward Attention for Expressive Speech Synthesis

    2021

    In speech interaction scenarios, speech emphasis is essential for expressing the underlying intention and attitude. Recently, end-to-end emphatic speech synthesis greatly improves the naturalness of synthetic speech, but also brings new problems: 1) lack of …

  5. Adversarially Learning Disentangled Speech Representations for Robust Multi-Factor Voice Conversion

    2021

    Factorizing speech as disentangled speech representations is vital to achieve highly controllable style transfer in voice conversion (VC).Conventional speech representation learning methods in VC only factorize speech as speaker and content, lacking controllability on other …

  6. Enhancing Speaking Styles in Conversational Text-to-Speech Synthesis with Graph-based Multi-modal Context Modeling

    2021 · arXiv (Cornell University)

    Comparing with traditional text-to-speech (TTS) systems, conversational TTS systems are required to synthesize speeches with proper speaking style confirming to the conversational context. However, state-of-the-art context modeling methods in conversational TTS only model the textual …

  7. An End-to-end Chinese Text Normalization Model based on Rule-guided Flat-Lattice Transformer

    2022 · arXiv (Cornell University)

    Text normalization, defined as a procedure transforming non standard words to spoken-form words, is crucial to the intelligibility of synthesized speech in text-to-speech system. Rule-based methods without considering context can not eliminate ambiguation, whereas sequence-to-sequence …

  8. Towards Multi-Scale Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech Synthesis

    2022 · arXiv (Cornell University)

    Previous works on expressive speech synthesis focus on modelling the mono-scale style embedding from the current sentence or context, but the multi-scale nature of speaking style in human speech is neglected. In this paper, we …

  9. Context-aware Coherent Speaking Style Prediction with Hierarchical Transformers for Audiobook Speech Synthesis

    2023 · arXiv (Cornell University)

    Recent advances in text-to-speech have significantly improved the expressiveness of synthesized speech. However, it is still challenging to generate speech with contextually appropriate and coherent speaking style for multi-sentence text in audiobooks. In this paper, …

  10. Enhancing the vocal range of single-speaker singing voice synthesis with melody-unsupervised pre-training

    2023 · arXiv (Cornell University)

    The single-speaker singing voice synthesis (SVS) usually underperforms at pitch values that are out of the singer's vocal range or associated with limited training samples. Based on our previous work, this work proposes a melody-unsupervised …

  11. VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling

    2024

    Recent AIGC systems possess the capability to generate digital multimedia content based on human language instructions, such as text, image and video. However, when it comes to speech, existing methods related to human instruction-to-speech generation …

  12. Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models

    2024 · arXiv (Cornell University)

    Spontaneous style speech synthesis, which aims to generate human-like speech, often encounters challenges due to the scarcity of high-quality data and limitations in model capabilities. Recent language model-based TTS systems can be trained on large, …

  13. StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion

    2025 · arXiv (Cornell University)

    Voice Conversion (VC) modifies speech to match a target speaker while preserving linguistic content. Traditional methods usually extract speaker information directly from speech while neglecting the explicit utilization of linguistic content. Since VC fundamentally involves …

  14. A Survey on In-context Learning

    2022 · arXiv (Cornell University)

    With the increasing capabilities of large language models (LLMs), in-context learning (ICL) has emerged as a new paradigm for natural language processing (NLP), where LLMs make predictions based on contexts augmented with a few examples. …