ملف الباحث

Frank K. Soong

5 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. KL-divergence based mispronunciation detection via DNN and decision tree in the phonetic space

    2016

    We propose to detect mispronunciations in a language learners speech via a discriminatively trained DNN in the phonetic space. The posterior probabilities of “senones” populated in a decision tree are trained and predicted speaker independently. …

  2. A Multi-Stage Multi-Codebook VQ-VAE Approach to High-Performance Neural TTS

    2022 · arXiv (Cornell University)

    We propose a Multi-Stage, Multi-Codebook (MSMC) approach to high-performance neural TTS synthesis. A vector-quantized, variational autoencoder (VQ-VAE) based feature analyzer is used to encode Mel spectrograms of speech training data by down-sampling progressively in multiple …

  3. Part-of-Speech Tagging with Bidirectional Long Short-Term Memory Recurrent Neural Network

    2015 · arXiv (Cornell University)

    Bidirectional Long Short-Term Memory Recurrent Neural Network (BLSTM-RNN) has been shown to be very effective for tagging sequential data, e.g. speech utterances or handwritten documents. While word embedding has been demoed as a powerful representation …

  4. A Unified Tagging Solution: Bidirectional LSTM Recurrent Neural Network with Word Embedding

    2015 · arXiv (Cornell University)

    Bidirectional Long Short-Term Memory Recurrent Neural Network (BLSTM-RNN) has been shown to be very effective for modeling and predicting sequential data, e.g. speech utterances or handwritten documents. In this study, we propose to use BLSTM-RNN …

  5. Speaker and language factorization in DNN-based TTS synthesis

    2016

    We have successfully proposed to use multi-speaker modelling in DNN-based TTS synthesis for improved voice quality with limited available data from a speaker. In this paper, we propose a new speaker and language factorized DNN, …