ملف الباحث

William Chan

6 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Deep Recurrent Neural Networks for Acoustic Modelling

    2015 · arXiv (Cornell University)

    We present a novel deep Recurrent Neural Network (RNN) model for acoustic modelling in Automatic Speech Recognition (ASR). We term our contribution as a TC-DNN-BLSTM-DNN model, the model combines a Deep Neural Network (DNN) with …

  2. End-to-End Speech Recognition Models

    2016 · KiltHub Repository

    For the past few decades, the bane of Automatic Speech Recognition (ASR) systems have been phonemes and Hidden Markov Models (HMMs). HMMs assume conditional indepen-dence between observations, and the reliance on explicit phonetic representations requires …

  3. Listen, attend and spell: A neural network for large vocabulary conversational speech recognition

    2016

    We present Listen, Attend and Spell (LAS), a neural speech recognizer that transcribes speech utterances directly to characters without pronunciation models, HMMs or other components of traditional speech recognizers. In LAS, the neural network architecture …

  4. Insertion Transformer: Flexible Sequence Generation via Insertion Operations

    2019 · arXiv (Cornell University)

    We present the Insertion Transformer, an iterative, partially autoregressive model for sequence generation based on insertion operations. Unlike typical autoregressive models which rely on a fixed, often left-to-right ordering of the output, our approach accommodates …

  5. Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling

    2019 · arXiv (Cornell University)

    Lingvo is a Tensorflow framework offering a complete solution for collaborative deep learning research, with a particular focus towards sequence-to-sequence models. Lingvo models are composed of modular building blocks that are flexible and easily extensible, …

  6. Bytes Are All You Need: End-to-end Multilingual Speech Recognition and Synthesis with Bytes

    2019

    We present two end-to-end models: Audio-to-Byte (A2B) and Byte-to-Audio (B2A), for multilingual speech recognition and synthesis. Prior work has predominantly used characters, sub-words or words as the unit of choice to model text. These units …