Jianhua Tao
9 papers in the PaperMetrix corpus
Papers by this author
-
Self-Attention Transducers for End-to-End Speech Recognition
2019
Recurrent neural network transducers (RNN-T) have been successfully applied in end-to-end speech recognition. However, the recurrent structure makes it difficult for parallelization . In this paper, we propose a self-attention transducer (SA-T) for speech recognition. …
-
Distilling Knowledge for Distant Speech Recognition via Parallel Data
2019
In order to improve the performance of distant speech recognition tasks, this paper proposes to distill knowledge from the close-talking model to the distant model using parallel data. The close-talking model is called the teacher …
-
One In A Hundred: Select The Best Predicted Sequence from Numerous Candidates for Streaming Speech Recognition
2020 · arXiv (Cornell University)
The RNN-Transducers and improved attention-based encoder-decoder models are widely applied to streaming speech recognition. Compared with these two end-to-end models, the CTC model is more efficient in training and inference. However, it cannot capture the …
-
Decoupling Pronunciation and Language for End-to-end Code-switching Automatic Speech Recognition
2020 · arXiv (Cornell University)
Despite the recent significant advances witnessed in end-to-end (E2E) ASR system for code-switching, hunger for audio-text paired data limits the further improvement of the models' performance. In this paper, we propose a decoupled transformer model …
-
TST: Time-Sparse Transducer for Automatic Speech Recognition
2023 · arXiv (Cornell University)
End-to-end model, especially Recurrent Neural Network Transducer (RNN-T), has achieved great success in speech recognition. However, transducer requires a great memory footprint and computing time when processing a long decoding sequence. To solve this problem, …
-
What Comes Next and Why? A Staged Encoder-Decoder Architecture for Script Event Prediction
2024
A script, which describes the evolutionary path of events, is a structured event sequence. Script event prediction aims to predict the next event from a sequence of historical events. Current studies favor modeling macroscale information, …
-
Transferring Personality Knowledge to Multimodal Sentiment Analysis
2024
Multimodal sentiment analysis systems have achieved remarkable success. However, significant challenges persist in tailoring sentiment analysis to individualized needs. Recognizing the pivotal role of personality traits in shaping emotional expression-where distinct personalities manifest emotions with …
-
DReSS: Data-driven Regularized Structured Streamlining for Large Language Models
2025 · arXiv (Cornell University)
Large language models (LLMs) have achieved significant progress across various domains, but their increasing scale results in high computational and memory costs. Recent studies have revealed that LLMs exhibit sparsity, providing the potential to reduce …
-
A Knowledge Distillation-Based Approach to Speech Emotion Recognition
2025 · IEEE Transactions on Affective Computing
Due to rapid advancements in deep learning, Transformer-based architectures have proven effective in speech emotion recognition (SER), largely due to their ability to model long-term dependencies more effectively than recurrent networks. The current Transformer architecture …