article

Deep Neural Network Approaches to Speaker and Language Recognition

  • IEEE Signal Processing Letters
  • Institute of Electrical and Electronics Engineers
Research footprint

At a glance

Citations
411
References
40
Comments
0
Paper overview

Öz

The impressive gains in performance obtained using deep neural networks (DNNs) for automatic speech recognition (ASR) have motivated the application of DNNs to other speech technologies such as speaker recognition (SR) and language recognition (LR). Prior work has shown performance gains for separate SR and LR tasks using DNNs for direct classification or for feature extraction. In this work we present the application of single DNN for both SR and LR using the 2013 Domain Adaptation Challenge speaker recognition (DAC13) and the NIST 2011 language recognition evaluation (LRE11) benchmarks. Using a single DNN trained for ASR on Switchboard data we demonstrate large gains on performance in both benchmarks: a 55% reduction in EER for the DAC13 out-of-domain condition and a 48% reduction in Cavg on the LRE11 30 s test condition. It is also shown that further gains are possible using score or feature fusion leading to the possibility of a single i-vector extractor producing state-of-the-art SR and LR performance.

Record transparency

Publication details

DOI
10.1109/lsp.2015.2420092
OpenAlex
W2078169166
Document type
article
Language
EN
Source
IEEE Signal Processing Letters
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.