Researcher profile

Yifan Gong

5 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. On Minimum Word Error Rate Training of the Hybrid Autoregressive Transducer

    2021

    Hybrid Autoregressive Transducer (HAT) is a recently proposed end-to-end acoustic model that extends the standard Recurrent Neural Network Transducer (RNN-T) for the purpose of the external language model (LM) fusion.In HAT, the blank probability and …

  2. Have Best of Both Worlds: Two-Pass Hybrid and E2E Cascading Framework for Speech Recognition

    2022 · ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

    Hybrid and end-to-end (E2E) systems have their individual advantages, with different error patterns in the speech recognition results. By jointly modeling audio and text, the E2E model performs better in matched scenarios and scales well …

  3. Have best of both worlds: two-pass hybrid and E2E cascading framework for speech recognition

    2021 · arXiv (Cornell University)

    Hybrid and end-to-end (E2E) systems have their individual advantages, with different error patterns in the speech recognition results. By jointly modeling audio and text, the E2E model performs better in matched scenarios and scales well …

  4. Streaming, fast and accurate on-device Inverse Text Normalization for Automatic Speech Recognition

    2022 · arXiv (Cornell University)

    Automatic Speech Recognition (ASR) systems typically yield output in lexical form. However, humans prefer a written form output. To bridge this gap, ASR systems usually employ Inverse Text Normalization (ITN). In previous works, Weighted Finite …

  5. End-to-End attention based text-dependent speaker verification

    2016

    A new type of End-to-End system for text-dependent speaker verification is presented in this paper. Previously, using the phonetic discriminate/speaker discriminate DNN as a feature extractor for speaker verification has shown promising results. The extracted …