Sanjeev Khudanpur
7 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Probing the Information Encoded in X-Vectors
2019 · 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
Deep neural network based speaker embeddings, such as x-vectors, have been shown to perform well in text-independent speaker recognition/verification tasks. In this paper, we use simple classifiers to investigate the contents encoded by x-vector embeddings. …
-
An Asynchronous WFST-Based Decoder For Automatic Speech Recognition
2021 · arXiv (Cornell University)
We introduce asynchronous dynamic decoder, which adopts an efficient A* algorithm to incorporate big language models in the one-pass decoding for large vocabulary continuous speech recognition. Unlike standard one-pass decoding with on-the-fly composition decoder which …
-
Adversarial Attacks and Defenses for Speech Recognition Systems
2021 · arXiv (Cornell University)
The ubiquitous presence of machine learning systems in our lives necessitates research into their vulnerabilities and appropriate countermeasures. In particular, we investigate the effectiveness of adversarial attacks and defenses against automatic speech recognition (ASR) systems. …
-
EURO: ESPnet Unsupervised ASR Open-source Toolkit
2022 · arXiv (Cornell University)
This paper describes the ESPnet Unsupervised ASR Open-source Toolkit (EURO), an end-to-end open-source toolkit for unsupervised automatic speech recognition (UASR). EURO adopts the state-of-the-art UASR learning method introduced by the Wav2vec-U, originally implemented at FAIRSEQ, …
-
Pronunciation and silence probability modeling for ASR
2015
In this paper we evaluate the WER improvement from modeling pronunciation probabilities and word-specific silence probabilities in speech recognition. We do this in the context of Finite State Transducer (FST)-based decoding, where pronunciation and silence …
-
A Pruned Rnnlm Lattice-Rescoring Algorithm for Automatic Speech Recognition
2018
Lattice-rescoring is a common approach to take advantage of recurrent neural language models in ASR, where a word-lattice is generated from 1st-pass decoding and the lattice is then rescored with a neural model, and ann-gram …
-
Semi-Supervised Training of Acoustic Models Using Lattice-Free MMI
2018
The lattice-free MMI objective (LF-MMI) has been used in supervised training of state-of-the-art neural network acoustic models for automatic speech recognition (ASR). With large amounts of unsupervised data available, extending this approach to the semi-supervised …