Researcher profile

Jinyu Li

9 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. On Minimum Word Error Rate Training of the Hybrid Autoregressive Transducer

    2021

    Hybrid Autoregressive Transducer (HAT) is a recently proposed end-to-end acoustic model that extends the standard Recurrent Neural Network Transducer (RNN-T) for the purpose of the external language model (LM) fusion.In HAT, the blank probability and …

  2. Have Best of Both Worlds: Two-Pass Hybrid and E2E Cascading Framework for Speech Recognition

    2022 · ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

    Hybrid and end-to-end (E2E) systems have their individual advantages, with different error patterns in the speech recognition results. By jointly modeling audio and text, the E2E model performs better in matched scenarios and scales well …

  3. Have best of both worlds: two-pass hybrid and E2E cascading framework for speech recognition

    2021 · arXiv (Cornell University)

    Hybrid and end-to-end (E2E) systems have their individual advantages, with different error patterns in the speech recognition results. By jointly modeling audio and text, the E2E model performs better in matched scenarios and scales well …

  4. Factorized Neural Transducer for Efficient Language Model Adaptation

    2021 · arXiv (Cornell University)

    In recent years, end-to-end (E2E) based automatic speech recognition (ASR) systems have achieved great success due to their simplicity and promising performance. Neural Transducer based models are increasingly popular in streaming E2E based ASR systems …

  5. CTCBERT: Advancing Hidden-unit BERT with CTC Objectives

    2022 · arXiv (Cornell University)

    In this work, we present a simple but effective method, CTCBERT, for advancing hidden-unit BERT (HuBERT). HuBERT applies a frame-level cross-entropy (CE) loss, which is similar to most acoustic model training. However, CTCBERT performs the …

  6. Joint Pre-Training with Speech and Bilingual Text for Direct Speech to Speech Translation

    2022 · arXiv (Cornell University)

    Direct speech-to-speech translation (S2ST) is an attractive research topic with many advantages compared to cascaded S2ST. However, direct S2ST suffers from the data scarcity problem because the corpora from speech of the source language to …

  7. Streaming, fast and accurate on-device Inverse Text Normalization for Automatic Speech Recognition

    2022 · arXiv (Cornell University)

    Automatic Speech Recognition (ASR) systems typically yield output in lexical form. However, humans prefer a written form output. To bridge this gap, ASR systems usually employ Inverse Text Normalization (ITN). In previous works, Weighted Finite …

  8. Accelerating Transducers through Adjacent Token Merging

    2023 · arXiv (Cornell University)

    Recent end-to-end automatic speech recognition (ASR) systems often utilize a Transformer-based acoustic encoder that generates embedding at a high frame rate. However, this design is inefficient, particularly for long speech signals due to the quadratic …

  9. End-to-End attention based text-dependent speaker verification

    2016

    A new type of End-to-End system for text-dependent speaker verification is presented in this paper. Previously, using the phonetic discriminate/speaker discriminate DNN as a feature extractor for speaker verification has shown promising results. The extracted …