ملف الباحث

Daniel Povey

6 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Probing the Information Encoded in X-Vectors

    2019 · 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)

    Deep neural network based speaker embeddings, such as x-vectors, have been shown to perform well in text-independent speaker recognition/verification tasks. In this paper, we use simple classifiers to investigate the contents encoded by x-vector embeddings. …

  2. An Asynchronous WFST-Based Decoder For Automatic Speech Recognition

    2021 · arXiv (Cornell University)

    We introduce asynchronous dynamic decoder, which adopts an efficient A* algorithm to incorporate big language models in the one-pass decoding for large vocabulary continuous speech recognition. Unlike standard one-pass decoding with on-the-fly composition decoder which …

  3. SUBLLM: A Novel Efficient Architecture with Token Sequence Subsampling for LLM

    2024 · arXiv (Cornell University)

    While Large Language Models (LLMs) have achieved remarkable success in various fields, the efficiency of training and inference remains a major challenge. To address this issue, we propose SUBLLM, short for Subsampling-Upsampling-Bypass Large Language Model, …

  4. Pronunciation and silence probability modeling for ASR

    2015

    In this paper we evaluate the WER improvement from modeling pronunciation probabilities and word-specific silence probabilities in speech recognition. We do this in the context of Finite State Transducer (FST)-based decoding, where pronunciation and silence …

  5. A Pruned Rnnlm Lattice-Rescoring Algorithm for Automatic Speech Recognition

    2018

    Lattice-rescoring is a common approach to take advantage of recurrent neural language models in ASR, where a word-lattice is generated from 1st-pass decoding and the lattice is then rescored with a neural model, and ann-gram …

  6. Semi-Supervised Training of Acoustic Models Using Lattice-Free MMI

    2018

    The lattice-free MMI objective (LF-MMI) has been used in supervised training of state-of-the-art neural network acoustic models for automatic speech recognition (ASR). With large amounts of unsupervised data available, extending this approach to the semi-supervised …