Yao Qian
6 papers in the PaperMetrix corpus
Papers by this author
-
A comparison of ASR and human errors for transcription of non-native spontaneous speech
2016
In this paper, we compare ASR and human transcriptions of non-native speech to investigate to what extent the accuracy and the patterns of errors of a modern ASR system match those of human listeners in …
-
Large-Scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification
2022 · ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
The speech representations learned from large-scale unlabeled data have shown better generalizability than those from supervised learning and thus attract a lot of interest to be applied for various downstream tasks. In this paper, we …
-
Deploying self-supervised learning in the wild for hybrid automatic speech recognition
2022 · arXiv (Cornell University)
Self-supervised learning (SSL) methods have proven to be very successful in automatic speech recognition (ASR). These great improvements have been reported mostly based on highly curated datasets such as LibriSpeech for non-streaming End-to-End ASR models. …
-
Part-of-Speech Tagging with Bidirectional Long Short-Term Memory Recurrent Neural Network
2015 · arXiv (Cornell University)
Bidirectional Long Short-Term Memory Recurrent Neural Network (BLSTM-RNN) has been shown to be very effective for tagging sequential data, e.g. speech utterances or handwritten documents. While word embedding has been demoed as a powerful representation …
-
A Unified Tagging Solution: Bidirectional LSTM Recurrent Neural Network with Word Embedding
2015 · arXiv (Cornell University)
Bidirectional Long Short-Term Memory Recurrent Neural Network (BLSTM-RNN) has been shown to be very effective for modeling and predicting sequential data, e.g. speech utterances or handwritten documents. In this study, we propose to use BLSTM-RNN …
-
Speaker and language factorization in DNN-based TTS synthesis
2016
We have successfully proposed to use multi-speaker modelling in DNN-based TTS synthesis for improved voice quality with limited available data from a speaker. In this paper, we propose a new speaker and language factorized DNN, …