Researcher profile

Chao Weng

5 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Towards Robust Speaker Verification with Target Speaker Enhancement

    2021 · arXiv (Cornell University)

    This paper proposes the target speaker enhancement based speaker verification network (TASE-SVNet), an all neural model that couples target speaker enhancement and speaker embedding extraction for robust speaker verification (SV). Specifically, an enrollment speaker conditioned …

  2. Towards end-to-end Speaker Diarization with Generalized Neural Speaker Clustering

    2022 · ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

    Speaker diarization consists of many components, e.g., front-end processing, speech activity detection (SAD), overlapped speech detection (OSD) and speaker segmentation/clustering. Conventionally, most of the involved components are separately developed and optimized. The resulting speaker diarization …

  3. Enhancing Speaking Styles in Conversational Text-to-Speech Synthesis with Graph-based Multi-modal Context Modeling

    2021 · arXiv (Cornell University)

    Comparing with traditional text-to-speech (TTS) systems, conversational TTS systems are required to synthesize speeches with proper speaking style confirming to the conversational context. However, state-of-the-art context modeling methods in conversational TTS only model the textual …

  4. Sifisinger: A High-Fidelity End-to-End Singing Voice Synthesizer Based on Source-Filter Model

    2024

    This paper presents an advanced end-to-end singing voice synthesis (SVS) system based on the source-filter mechanism that directly translates lyrical and melodic cues into expressive and high-fidelity human-like singing. Similarly to VISinger 2, the proposed …

  5. Investigating End-to-end Speech Recognition for Mandarin-english Code-switching

    2019

    Code-switching is a common phenomenon in many multilingual communities and presents a challenge to automatic speech recognition (ASR). In this paper, three approaches are investigated to improve end-to-end speech recognition on Mandarin-English code-switching task. First, …