Shujie Liu
7 papers in the PaperMetrix corpus
Papers by this author
-
Large-Scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification
2022 · ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
The speech representations learned from large-scale unlabeled data have shown better generalizability than those from supervised learning and thus attract a lot of interest to be applied for various downstream tasks. In this paper, we …
-
Joint Pre-Training with Speech and Bilingual Text for Direct Speech to Speech Translation
2022 · arXiv (Cornell University)
Direct speech-to-speech translation (S2ST) is an attractive research topic with many advantages compared to cascaded S2ST. However, direct S2ST suffers from the data scarcity problem because the corpora from speech of the source language to …
-
Accelerating Transducers through Adjacent Token Merging
2023 · arXiv (Cornell University)
Recent end-to-end automatic speech recognition (ASR) systems often utilize a Transformer-based acoustic encoder that generates embedding at a high frame rate. However, this design is inefficient, particularly for long speech signals due to the quadratic …
-
MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning
2025 · arXiv (Cornell University)
Medical Large Vision-Language Models (Med-LVLMs) have shown strong potential in multimodal diagnostic tasks. However, existing single-agent models struggle to generalize across diverse medical specialties, limiting their performance. Recent efforts introduce multi-agent collaboration frameworks inspired by …
-
Achieving Human Parity on Automatic Chinese to English News Translation
2018 · arXiv (Cornell University)
Machine translation has made rapid advances in recent years. Millions of people are using it today in online translation systems and mobile applications in order to communicate across language barriers. The question naturally arises whether …
-
Style Transfer as Unsupervised Machine Translation
2018 · arXiv (Cornell University)
Language style transferring rephrases text with specific stylistic attributes while preserving the original attribute-independent content. One main challenge in learning a style transfer system is a lack of parallel data where the source sentence is …
-
Neural Speech Synthesis with Transformer Network
2019 · Proceedings of the AAAI Conference on Artificial Intelligence
Although end-to-end neural text-to-speech (TTS) methods (such as Tacotron2) are proposed and achieve state-of-theart performance, they still suffer from two problems: 1) low efficiency during training and inference; 2) hard to model long dependency using …