Kanishka Rao
7 papers in the PaperMetrix corpus
Papers by this author
-
Automatic pronunciation verification for speech recognition
2015
Pronunciations for words are a critical component in an automated speech recognition system (ASR) as mis-recognitions may be caused by missing or inaccurate pronunciations. The need for high quality pronunciations has recently motivated data-driven techniques …
-
Manner of Articulation based Split Lattices for Phoneme Recognition
2018
Phoneme lattices have been shown to be a good choice to encode in a compact way alternative decoding hypotheses from a speech recognition system. However the optimal phoneme sequence is produced by tracing all the …
-
Grapheme-to-phoneme conversion using Long Short-Term Memory recurrent neural networks
2015
Grapheme-to-phoneme (G2P) models are key components in speech recognition and text-to-speech systems as they describe how words are pronounced. We propose a G2P model based on a Long Short-Term Memory (LSTM) recurrent neural network (RNN). …
-
Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling
2019 · arXiv (Cornell University)
Lingvo is a Tensorflow framework offering a complete solution for collaborative deep learning research, with a particular focus towards sequence-to-sequence models. Lingvo models are composed of modular building blocks that are flexible and easily extensible, …
-
Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-transducer
2017 · 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
We investigate training end-to-end speech recognition models with the recurrent neural network transducer (RNN-T): a streaming, all-neural, sequence-to-sequence architecture which jointly learns acoustic and language model components from transcribed acoustic data. We explore various model …
-
Multilingual Speech Recognition with a Single End-to-End Model
2018
Training a conventional automatic speech recognition (ASR) system to support multiple languages is challenging because the sub-word unit, lexicon and word inventories are typically language specific. In contrast, sequence-to-sequence models are well suited for multilingual …
-
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
2022 · arXiv (Cornell University)
Large language models can encode a wealth of semantic knowledge about the world. Such knowledge could be extremely useful to robots aiming to act upon high-level, temporally extended instructions expressed in natural language. However, a …