Ziyang Ma
4 papers in the PaperMetrix corpus
Papers by this author
-
ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering
2024 · arXiv (Cornell University)
The language model (LM) approach based on acoustic and linguistic prompts, such as VALL-E, has achieved remarkable progress in the field of zero-shot audio generation. However, existing methods still have some limitations: 1) repetitions, transpositions, …
-
TacoLM: GaTed Attention Equipped Codec Language Model are Efficient Zero-Shot Text to Speech Synthesizers
2024 · arXiv (Cornell University)
Neural codec language model (LM) has demonstrated strong capability in zero-shot text-to-speech (TTS) synthesis. However, the codec LM often suffers from limitations in inference speed and stability, due to its auto-regressive nature and implicit alignment …
-
CTC-Assisted LLM-Based Contextual ASR
2024
Contextual ASR or hotword customization holds substantial practical value. Despite the impressive performance of current end-to-end (E2E) automatic speech recognition (ASR) systems, they often face challenges in accurately recognizing rare words. Typical E2E contextual ASR …
-
Towards Flow-Matching-based TTS without Classifier-Free Guidance
2025 · arXiv (Cornell University)
Flow matching has demonstrated strong generative capabilities and has become a core component in modern Text-to-Speech (TTS) systems. To ensure high-quality speech synthesis, Classifier-Free Guidance (CFG) is widely used during the inference of flow-matching-based TTS …