Yifan Gong
5 papers in the PaperMetrix corpus
Papers by this author
-
On Minimum Word Error Rate Training of the Hybrid Autoregressive Transducer
2021
Hybrid Autoregressive Transducer (HAT) is a recently proposed end-to-end acoustic model that extends the standard Recurrent Neural Network Transducer (RNN-T) for the purpose of the external language model (LM) fusion.In HAT, the blank probability and …
-
Have Best of Both Worlds: Two-Pass Hybrid and E2E Cascading Framework for Speech Recognition
2022 · ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Hybrid and end-to-end (E2E) systems have their individual advantages, with different error patterns in the speech recognition results. By jointly modeling audio and text, the E2E model performs better in matched scenarios and scales well …
-
Have best of both worlds: two-pass hybrid and E2E cascading framework for speech recognition
2021 · arXiv (Cornell University)
Hybrid and end-to-end (E2E) systems have their individual advantages, with different error patterns in the speech recognition results. By jointly modeling audio and text, the E2E model performs better in matched scenarios and scales well …
-
Streaming, fast and accurate on-device Inverse Text Normalization for Automatic Speech Recognition
2022 · arXiv (Cornell University)
Automatic Speech Recognition (ASR) systems typically yield output in lexical form. However, humans prefer a written form output. To bridge this gap, ASR systems usually employ Inverse Text Normalization (ITN). In previous works, Weighted Finite …
-
End-to-End attention based text-dependent speaker verification
2016
A new type of End-to-End system for text-dependent speaker verification is presented in this paper. Previously, using the phonetic discriminate/speaker discriminate DNN as a feature extractor for speaker verification has shown promising results. The extracted …