Improving Unsupervised Sparsespeech Acoustic Models with Categorical\n Reparameterization
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Abstract
The Sparsespeech model is an unsupervised acoustic model that can generate\ndiscrete pseudo-labels for untranscribed speech. We extend the Sparsespeech\nmodel to allow for sampling over a random discrete variable, yielding\npseudo-posteriorgrams. The degree of sparsity in this posteriorgram can be\nfully controlled after the model has been trained. We use the Gumbel-Softmax\ntrick to approximately sample from a discrete distribution in the neural\nnetwork and this allows us to train the network efficiently with standard\nbackpropagation. The new and improved model is trained and evaluated on the\nLibri-Light corpus, a benchmark for ASR with limited or no supervision. The\nmodel is trained on 600h and 6000h of English read speech. We evaluate the\nimproved model using the ABX error measure and a semi-supervised setting with\n10h of transcribed speech. We observe a relative improvement of up to 31.4% on\nABX error rates across speakers on the test set with the improved Sparsespeech\nmodel on 600h of speech data and further improvements when we scale the model\nto 6000h.\n
Publication details
- DOI
- 10.48550/arxiv.2005.14578
- OpenAlex
- W4287774020
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
Log in to join the discussion.