Researcher profile

Xixin Wu

7 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Recurrent Neural Network Language Model Training Using Natural Gradient

    2019

    Recurrent neural network language models (RNNLMs) have become an increasing popular choice for state-of-the-art speech recognition systems. RNNLMs are normally trained by minimizing the cross entropy (CE) using the stochastic gradient descent (SGD) algorithm. However, …

  2. Improved End-to-End Dysarthric Speech Recognition via Meta-learning Based Model Re-initialization

    2021

    Dysarthric speech recognition is a challenging task as dysarthric data is limited and its acoustics deviate significantly from normal speech. Model-based speaker adaptation is a promising method by using the limited dysarthric speech to fine-tune …

  3. Channel-wise Gated Res2Net: Towards Robust Detection of Synthetic Speech Attacks

    2021 · arXiv (Cornell University)

    Existing approaches for anti-spoofing in automatic speaker verification (ASV) still lack generalizability to unseen attacks. The Res2Net approach designs a residual-like connection between feature groups within one block, which increases the possible receptive fields and …

  4. A Multi-Stage Multi-Codebook VQ-VAE Approach to High-Performance Neural TTS

    2022 · arXiv (Cornell University)

    We propose a Multi-Stage, Multi-Codebook (MSMC) approach to high-performance neural TTS synthesis. A vector-quantized, variational autoencoder (VQ-VAE) based feature analyzer is used to encode Mel spectrograms of speech training data by down-sampling progressively in multiple …

  5. A Multi-Scale Time-Frequency Spectrogram Discriminator for GAN-based Non-Autoregressive TTS

    2022 · Interspeech 2022

    The generative adversarial network (GAN) has shown its outstanding capability in improving Non-Autoregressive TTS (NAR-TTS) by adversarially training it with an extra model that discriminates between the real and the generated speech.To maximize the benefits …

  6. Seamless Language Expansion: Enhancing Multilingual Mastery in Self-Supervised Models

    2024 · arXiv (Cornell University)

    Self-supervised (SSL) models have shown great performance in various downstream tasks. However, they are typically developed for limited languages, and may encounter new languages in real-world. Developing a SSL model for each new language is …

  7. Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models

    2024 · arXiv (Cornell University)

    Spontaneous style speech synthesis, which aims to generate human-like speech, often encounters challenges due to the scarcity of high-quality data and limitations in model capabilities. Recent language model-based TTS systems can be trained on large, …