preprint Open access

Improving Unsupervised Sparsespeech Acoustic Models with Categorical\n Reparameterization

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Abstract

The Sparsespeech model is an unsupervised acoustic model that can generate\ndiscrete pseudo-labels for untranscribed speech. We extend the Sparsespeech\nmodel to allow for sampling over a random discrete variable, yielding\npseudo-posteriorgrams. The degree of sparsity in this posteriorgram can be\nfully controlled after the model has been trained. We use the Gumbel-Softmax\ntrick to approximately sample from a discrete distribution in the neural\nnetwork and this allows us to train the network efficiently with standard\nbackpropagation. The new and improved model is trained and evaluated on the\nLibri-Light corpus, a benchmark for ASR with limited or no supervision. The\nmodel is trained on 600h and 6000h of English read speech. We evaluate the\nimproved model using the ABX error measure and a semi-supervised setting with\n10h of transcribed speech. We observe a relative improvement of up to 31.4% on\nABX error rates across speakers on the test set with the improved Sparsespeech\nmodel on 600h of speech data and further improvements when we scale the model\nto 6000h.\n

Record transparency

Publication details

DOI
10.48550/arxiv.2005.14578
OpenAlex
W4287774020
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.