Researcher profile

Kanishka Rao

7 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Automatic pronunciation verification for speech recognition

    2015

    Pronunciations for words are a critical component in an automated speech recognition system (ASR) as mis-recognitions may be caused by missing or inaccurate pronunciations. The need for high quality pronunciations has recently motivated data-driven techniques …

  2. Manner of Articulation based Split Lattices for Phoneme Recognition

    2018

    Phoneme lattices have been shown to be a good choice to encode in a compact way alternative decoding hypotheses from a speech recognition system. However the optimal phoneme sequence is produced by tracing all the …

  3. Grapheme-to-phoneme conversion using Long Short-Term Memory recurrent neural networks

    2015

    Grapheme-to-phoneme (G2P) models are key components in speech recognition and text-to-speech systems as they describe how words are pronounced. We propose a G2P model based on a Long Short-Term Memory (LSTM) recurrent neural network (RNN). …

  4. Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling

    2019 · arXiv (Cornell University)

    Lingvo is a Tensorflow framework offering a complete solution for collaborative deep learning research, with a particular focus towards sequence-to-sequence models. Lingvo models are composed of modular building blocks that are flexible and easily extensible, …

  5. Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-transducer

    2017 · 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)

    We investigate training end-to-end speech recognition models with the recurrent neural network transducer (RNN-T): a streaming, all-neural, sequence-to-sequence architecture which jointly learns acoustic and language model components from transcribed acoustic data. We explore various model …

  6. Multilingual Speech Recognition with a Single End-to-End Model

    2018

    Training a conventional automatic speech recognition (ASR) system to support multiple languages is challenging because the sub-word unit, lexicon and word inventories are typically language specific. In contrast, sequence-to-sequence models are well suited for multilingual …

  7. Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

    2022 · arXiv (Cornell University)

    Large language models can encode a wealth of semantic knowledge about the world. Such knowledge could be extremely useful to robots aiming to act upon high-level, temporally extended instructions expressed in natural language. However, a …