Haşim Sak
4 papers in the PaperMetrix corpus
Papers by this author
-
Clustering and Mining Accented Speech for Inclusive and Fair Speech Recognition
2024 · arXiv (Cornell University)
Modern automatic speech recognition (ASR) systems are typically trained on more than tens of thousands hours of speech data, which is one of the main factors for their great success. However, the distribution of such …
-
Unidirectional long short-term memory recurrent neural network with recurrent output layer for low-latency speech synthesis
2015
Long short-term memory recurrent neural networks (LSTM-RNNs) have been applied to various speech applications including acoustic modeling for statistical parametric speech synthesis. One of the concerns for applying them to text-to-speech applications is its effect …
-
Grapheme-to-phoneme conversion using Long Short-Term Memory recurrent neural networks
2015
Grapheme-to-phoneme (G2P) models are key components in speech recognition and text-to-speech systems as they describe how words are pronounced. We propose a G2P model based on a Long Short-Term Memory (LSTM) recurrent neural network (RNN). …
-
Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-transducer
2017 · 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
We investigate training end-to-end speech recognition models with the recurrent neural network transducer (RNN-T): a streaming, all-neural, sequence-to-sequence architecture which jointly learns acoustic and language model components from transcribed acoustic data. We explore various model …