Xuankai Chang
6 papers in the PaperMetrix corpus
Papers by this author
-
Recognizing Multi-talker Speech with Permutation Invariant Training
2017 · arXiv (Cornell University)
In this paper, we propose a novel technique for direct recognition of multiple speech streams given the single channel of mixed speech, without first separating them. Our technique is based on permutation invariant training (PIT) …
-
Monaural Multi-Talker Speech Recognition with Attention Mechanism and Gated Convolutional Networks
2018
Provided are a speech recognition training processing method and an apparatus including the same. The speech recognition training processing method includes acquiring multi-talker mixed speech sequence data corresponding to a plurality of speakers, encoding the …
-
Streaming End-to-End ASR based on Blockwise Non-Autoregressive Models
2021 · arXiv (Cornell University)
Non-autoregressive (NAR) modeling has gained more and more attention in speech processing. With recent state-of-the-art attention-based automatic speech recognition (ASR) structure, NAR can realize promising real-time factor (RTF) improvement with only small degradation of accuracy …
-
End-to-End Integration of Speech Recognition, Speech Enhancement, and Self-Supervised Learning Representation
2022 · arXiv (Cornell University)
This work presents our end-to-end (E2E) automatic speech recognition (ASR) model targetting at robust speech recognition, called Integraded speech Recognition with enhanced speech Input for Self-supervised learning representation (IRIS). Compared with conventional E2E ASR models, …
-
SynesLM: A Unified Approach for Audio-visual Speech Recognition and Translation via Language Model and Synthetic Data
2024 · arXiv (Cornell University)
In this work, we present SynesLM, an unified model which can perform three multimodal language understanding tasks: audio-visual automatic speech recognition(AV-ASR) and visual-aided speech/machine translation(VST/VMT). Unlike previous research that focused on lip motion as visual …
-
SUPERB: Speech Processing Universal PERformance Benchmark
2021
Self-supervised learning (SSL) has proven vital for advancing research in natural language processing (NLP) and computer vision (CV).The paradigm pretrains a shared model on large volumes of unlabeled data and achieves state-of-the-art (SOTA) for various …