Thomas Fang Zheng
3 papers in the PaperMetrix corpus
Papers by this author
-
Neural Discriminant Analysis for Deep Speaker Embedding
2020 · arXiv (Cornell University)
Probabilistic Linear Discriminant Analysis (PLDA) is a popular tool in open-set classification/verification tasks. However, the Gaussian assumption underlying PLDA prevents it from being applied to situations where the data is clearly non-Gaussian. In this paper, …
-
Speaker Adaptation for Quantised End-to-End ASR Models
2024 · arXiv (Cornell University)
End-to-end models have shown superior performance for automatic speech recognition (ASR). However, such models are often very large in size and thus challenging to deploy on resource-constrained edge devices. While quantisation can reduce model sizes, …
-
Transfer learning for speech and language processing
2015
Transfer learning is a vital technique that generalizes models trained for one setting or task to other settings or tasks. For example in speech recognition, an acoustic model trained for one language can be used …