conference-paper Open access

A Temporal Coherence Loss Function for Learning Unsupervised Acoustic Embeddings

  • Procedia Computer Science
  • Elsevier BV
Research footprint

At a glance

Citations
20
References
26
Comments
0
Paper overview

Öz

We train neural networks of varying depth with a loss function which imposes the output representations to have a temporal profile which looks like that of phonemes. We show that a simple loss function which maximizes the dissimilarity between near frames and long distance frames helps to construct a speech embedding that improves phoneme discriminability, both within and across speakers, even though the loss function only uses within speaker information. However, with too deep an architecture, this loss function yields overfitting, suggesting the need for more data and/or regularization.

Record transparency

Publication details

DOI
10.1016/j.procs.2016.04.035
OpenAlex
W2345968833
Document type
conference-paper
Language
EN
Source
Procedia Computer Science
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.