Researcher profile

Xie Chen

9 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Recurrent Neural Network Language Model Training Using Natural Gradient

    2019

    Recurrent neural network language models (RNNLMs) have become an increasing popular choice for state-of-the-art speech recognition systems. RNNLMs are normally trained by minimizing the cross entropy (CE) using the stochastic gradient descent (SGD) algorithm. However, …

  2. Information scrambling in chaotic systems with dissipation

    2018 · arXiv (Cornell University)

    Chaotic dynamics in closed local quantum systems scrambles quantum information, which is manifested quantitatively in the decay of the out-of-time-ordered correlators (OTOC) of local operators. How is information scrambling affected when the system is coupled …

  3. Factorized Neural Transducer for Efficient Language Model Adaptation

    2021 · arXiv (Cornell University)

    In recent years, end-to-end (E2E) based automatic speech recognition (ASR) systems have achieved great success due to their simplicity and promising performance. Neural Transducer based models are increasingly popular in streaming E2E based ASR systems …

  4. Improved Factorized Neural Transducer Model For text-only Domain Adaptation

    2023 · arXiv (Cornell University)

    Adapting End-to-End ASR models to out-of-domain datasets with text data is challenging. Factorized neural Transducer (FNT) aims to address this issue by introducing a separate vocabulary decoder to predict the vocabulary. Nonetheless, this approach has …

  5. ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

    2024 · arXiv (Cornell University)

    The language model (LM) approach based on acoustic and linguistic prompts, such as VALL-E, has achieved remarkable progress in the field of zero-shot audio generation. However, existing methods still have some limitations: 1) repetitions, transpositions, …

  6. StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations

    2024

    While acoustic expressiveness has long been studied in expressive text-to-speech (ETTS), the inherent expressiveness in text lacks sufficient attention, especially for ETTS of artistic works. In this paper, we introduce StoryTTS, a highly ETTS dataset …

  7. TacoLM: GaTed Attention Equipped Codec Language Model are Efficient Zero-Shot Text to Speech Synthesizers

    2024 · arXiv (Cornell University)

    Neural codec language model (LM) has demonstrated strong capability in zero-shot text-to-speech (TTS) synthesis. However, the codec LM often suffers from limitations in inference speed and stability, due to its auto-regressive nature and implicit alignment …

  8. CTC-Assisted LLM-Based Contextual ASR

    2024

    Contextual ASR or hotword customization holds substantial practical value. Despite the impressive performance of current end-to-end (E2E) automatic speech recognition (ASR) systems, they often face challenges in accurately recognizing rare words. Typical E2E contextual ASR …

  9. Towards Flow-Matching-based TTS without Classifier-Free Guidance

    2025 · arXiv (Cornell University)

    Flow matching has demonstrated strong generative capabilities and has become a core component in modern Text-to-Speech (TTS) systems. To ensure high-quality speech synthesis, Classifier-Free Guidance (CFG) is widely used during the inference of flow-matching-based TTS …