Li-Rong Dai
5 papers in the PaperMetrix corpus
Papers by this author
-
Joint training of front-end and back-end deep neural networks for robust speech recognition
2015
Based on the recently proposed speech pre-processing front-end with deep neural networks (DNNs), we first investigate different feature mapping directly from noisy speech via DNN for robust speech recognition. Next, we propose to jointly train …
-
State-Clustering Based Multiple Deep Neural Networks Modeling Approach for Speech Recognition
2015 · IEEE/ACM Transactions on Audio Speech and Language Processing
The hybrid deep neural network (DNN) and hidden Markov model (HMM) has recently achieved dramatic performance gains in automatic speech recognition (ASR). The DNN-based acoustic model is very powerful but its learning process is extremely …
-
Unsupervised speaker adaptation of BLSTM-RNN for LVCSR based on speaker code
2016
Recently, the speaker code based adaptation has been successfully expanded to recurrent neural networks using bidirectional Long Short-Term Memory (BLSTM-RNN) [1]. Experiments on the small-scale TIMIT task have demonstrated that the speaker code based adaptation …
-
Sifisinger: A High-Fidelity End-to-End Singing Voice Synthesizer Based on Source-Filter Model
2024
This paper presents an advanced end-to-end singing voice synthesis (SVS) system based on the source-filter mechanism that directly translates lyrical and melodic cues into expressive and high-fidelity human-like singing. Similarly to VISinger 2, the proposed …
-
Forward Attention in Sequence- To-Sequence Acoustic Modeling for Speech Synthesis
2018
This paper proposes a forward attention method for the sequence-to-sequence acoustic modeling of speech synthesis. This method is motivated by the nature of the monotonic alignment from phone sequences to acoustic sequences. Only the alignment …