Researcher profile

Haşim Sak

4 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Clustering and Mining Accented Speech for Inclusive and Fair Speech Recognition

    2024 · arXiv (Cornell University)

    Modern automatic speech recognition (ASR) systems are typically trained on more than tens of thousands hours of speech data, which is one of the main factors for their great success. However, the distribution of such …

  2. Unidirectional long short-term memory recurrent neural network with recurrent output layer for low-latency speech synthesis

    2015

    Long short-term memory recurrent neural networks (LSTM-RNNs) have been applied to various speech applications including acoustic modeling for statistical parametric speech synthesis. One of the concerns for applying them to text-to-speech applications is its effect …

  3. Grapheme-to-phoneme conversion using Long Short-Term Memory recurrent neural networks

    2015

    Grapheme-to-phoneme (G2P) models are key components in speech recognition and text-to-speech systems as they describe how words are pronounced. We propose a G2P model based on a Long Short-Term Memory (LSTM) recurrent neural network (RNN). …

  4. Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-transducer

    2017 · 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)

    We investigate training end-to-end speech recognition models with the recurrent neural network transducer (RNN-T): a streaming, all-neural, sequence-to-sequence architecture which jointly learns acoustic and language model components from transcribed acoustic data. We explore various model …