Xixin Wu
7 papers in the PaperMetrix corpus
Papers by this author
-
Recurrent Neural Network Language Model Training Using Natural Gradient
2019
Recurrent neural network language models (RNNLMs) have become an increasing popular choice for state-of-the-art speech recognition systems. RNNLMs are normally trained by minimizing the cross entropy (CE) using the stochastic gradient descent (SGD) algorithm. However, …
-
Improved End-to-End Dysarthric Speech Recognition via Meta-learning Based Model Re-initialization
2021
Dysarthric speech recognition is a challenging task as dysarthric data is limited and its acoustics deviate significantly from normal speech. Model-based speaker adaptation is a promising method by using the limited dysarthric speech to fine-tune …
-
Channel-wise Gated Res2Net: Towards Robust Detection of Synthetic Speech Attacks
2021 · arXiv (Cornell University)
Existing approaches for anti-spoofing in automatic speaker verification (ASV) still lack generalizability to unseen attacks. The Res2Net approach designs a residual-like connection between feature groups within one block, which increases the possible receptive fields and …
-
A Multi-Stage Multi-Codebook VQ-VAE Approach to High-Performance Neural TTS
2022 · arXiv (Cornell University)
We propose a Multi-Stage, Multi-Codebook (MSMC) approach to high-performance neural TTS synthesis. A vector-quantized, variational autoencoder (VQ-VAE) based feature analyzer is used to encode Mel spectrograms of speech training data by down-sampling progressively in multiple …
-
A Multi-Scale Time-Frequency Spectrogram Discriminator for GAN-based Non-Autoregressive TTS
2022 · Interspeech 2022
The generative adversarial network (GAN) has shown its outstanding capability in improving Non-Autoregressive TTS (NAR-TTS) by adversarially training it with an extra model that discriminates between the real and the generated speech.To maximize the benefits …
-
Seamless Language Expansion: Enhancing Multilingual Mastery in Self-Supervised Models
2024 · arXiv (Cornell University)
Self-supervised (SSL) models have shown great performance in various downstream tasks. However, they are typically developed for limited languages, and may encounter new languages in real-world. Developing a SSL model for each new language is …
-
Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
2024 · arXiv (Cornell University)
Spontaneous style speech synthesis, which aims to generate human-like speech, often encounters challenges due to the scarcity of high-quality data and limitations in model capabilities. Recent language model-based TTS systems can be trained on large, …