Frank K. Soong
5 papers in the PaperMetrix corpus
Papers by this author
-
KL-divergence based mispronunciation detection via DNN and decision tree in the phonetic space
2016
We propose to detect mispronunciations in a language learners speech via a discriminatively trained DNN in the phonetic space. The posterior probabilities of “senones” populated in a decision tree are trained and predicted speaker independently. …
-
A Multi-Stage Multi-Codebook VQ-VAE Approach to High-Performance Neural TTS
2022 · arXiv (Cornell University)
We propose a Multi-Stage, Multi-Codebook (MSMC) approach to high-performance neural TTS synthesis. A vector-quantized, variational autoencoder (VQ-VAE) based feature analyzer is used to encode Mel spectrograms of speech training data by down-sampling progressively in multiple …
-
Part-of-Speech Tagging with Bidirectional Long Short-Term Memory Recurrent Neural Network
2015 · arXiv (Cornell University)
Bidirectional Long Short-Term Memory Recurrent Neural Network (BLSTM-RNN) has been shown to be very effective for tagging sequential data, e.g. speech utterances or handwritten documents. While word embedding has been demoed as a powerful representation …
-
A Unified Tagging Solution: Bidirectional LSTM Recurrent Neural Network with Word Embedding
2015 · arXiv (Cornell University)
Bidirectional Long Short-Term Memory Recurrent Neural Network (BLSTM-RNN) has been shown to be very effective for modeling and predicting sequential data, e.g. speech utterances or handwritten documents. In this study, we propose to use BLSTM-RNN …
-
Speaker and language factorization in DNN-based TTS synthesis
2016
We have successfully proposed to use multi-speaker modelling in DNN-based TTS synthesis for improved voice quality with limited available data from a speaker. In this paper, we propose a new speaker and language factorized DNN, …