conference-paper

Discriminative Deep Audio Feature Embedding for Speaker Recognition in the Wild

Research footprint

At a glance

Citations
4
References
25
Comments
0
Paper overview

Öz

In this paper we face the problem of speaker recognition in the wild. We tackle the speaker identification and verification problems with the use of Deep Convolutional Neural Networks (CNN). We propose the modification of two Residual CNN architectures (ResNet) in order to be used with the spectrograms of the audio data as input images. The proposed architectures, trained with a contrastive-center loss, have been tested on the VoxCeleb and SIWIS datasets on both the speaker identification and verification tasks. The experimental results show the effectiveness of the proposed solution with respect to the state of the art. The proposed network shows to be robust in unconstrained conditions and, more important, it shows to be quite robust in a multilingual scenario.

Record transparency

Publication details

DOI
10.1109/icce-berlin.2018.8576237
OpenAlex
W2904230401
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.