Xiaodong Shi
4 papers in the PaperMetrix corpus
Papers by this author
-
Improving Chinese Pop Song and Hokkien Gezi Opera Singing Voice Synthesis by Enhancing Local Modeling
2023
Singing Voice Synthesis (SVS) strives to synthesize pleasing vocals based on music scores and lyrics. The current acoustic models based on Transformer usually process the entire sequence globally and use a simple L1 loss. However, …
-
UniMem: Towards a Unified View of Long-Context Large Language Models
2024 · arXiv (Cornell University)
Long-context processing is a critical ability that constrains the applicability of large language models (LLMs). Although there exist various methods devoted to enhancing the long-context processing ability of LLMs, they are developed in an isolated …
-
Representation Purification for End-to-End Speech Translation
2024 · arXiv (Cornell University)
Speech-to-text translation (ST) is a cross-modal task that involves converting spoken language into text in a different language. Previous research primarily focused on enhancing speech translation by facilitating knowledge transfer from machine translation, exploring various …
-
Deep Semantic Role Labeling With Self-Attention
2018 · Proceedings of the AAAI Conference on Artificial Intelligence
Semantic Role Labeling (SRL) is believed to be a crucial step towards natural language understanding and has been widely studied. Recent years, end-to-end SRL with recurrent neural networks (RNN) has gained increasing attention. However, it …