conference-paper

KL-divergence based mispronunciation detection via DNN and decision tree in the phonetic space

Research footprint

At a glance

Citations
3
References
16
Comments
0
Paper overview

Abstract

We propose to detect mispronunciations in a language learners speech via a discriminatively trained DNN in the phonetic space. The posterior probabilities of “senones” populated in a decision tree are trained and predicted speaker independently. Acoustic features of each input segment (with preceding and succeeding contexts of several frames) are mapped unto the whole set of senones in their corresponding posteriors. Vectors of senone posteriors are used as stochastic characterization of input speech segments in the phonetic space. Distortion between any two such vectors are measured with the symmetric Kullback-Leibler Divergence (KLD) and they are used for performing vector clustering and computing the corresponding centroids in a phonetically oriented senone based decision tree. Experimental results, tested on a large, Mandarin database (iCALL) of L2 language learners, show that the proposed approach to mispronunciation detection can achieve a 3.0% of equal precision and recall improvement over our best DNN-based, Goodness of Pronunciation (GOP) baseline system with adaptation. When the original maximum-likelihood trained decision tree is retrained with the symmetric KLD measure, further improvement of 0.8% of equal precision and recall can be obtained.

Record transparency

Publication details

DOI
10.1109/apsipa.2016.7820849
OpenAlex
W2578750458
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.