conference-paper
Spoken Language Identification with Deep Convolutional Neural Network and Data Augmentation
Research footprint
At a glance
- Citations
- 1
- References
- 2
- Comments
- 0
Paper overview
Abstract
In this paper, a spoken language detection system based on deep convolutional neural networks is presented. The neural network model is trained and tested on a speech dataset containing five languages. Speech signals are first converted into mel-spectrogram features and these features are fed into the deep convolutional neural network. Flattened outputs of the deep convolutional network are then fed into a recurrent layer, and a dense layer with softmax activation function is used as an output layer to predict the output language probabilities. This network results in 0.89 Fl-score in our test data. We also used a data augmentation method, namely SpecAugment, which increased the Fl-score to 0.94.
Record transparency
Publication details
- DOI
- 10.1109/siu49456.2020.9302425
- OpenAlex
- W3134309788
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.