conference-paper

Spoken Language Identification with Deep Convolutional Neural Network and Data Augmentation

Research footprint

At a glance

Citations
1
References
2
Comments
0
Paper overview

Abstract

In this paper, a spoken language detection system based on deep convolutional neural networks is presented. The neural network model is trained and tested on a speech dataset containing five languages. Speech signals are first converted into mel-spectrogram features and these features are fed into the deep convolutional neural network. Flattened outputs of the deep convolutional network are then fed into a recurrent layer, and a dense layer with softmax activation function is used as an output layer to predict the output language probabilities. This network results in 0.89 Fl-score in our test data. We also used a data augmentation method, namely SpecAugment, which increased the Fl-score to 0.94.

Record transparency

Publication details

DOI
10.1109/siu49456.2020.9302425
OpenAlex
W3134309788
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.