preprint Open access

vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
169
References
34
Comments
0
Paper overview

Abstract

We propose vq-wav2vec to learn discrete representations of audio segments through a wav2vec-style self-supervised context prediction task. The algorithm uses either a gumbel softmax or online k-means clustering to quantize the dense representations. Discretization enables the direct application of algorithms from the NLP community which require discrete inputs. Experiments show that BERT pre-training achieves a new state of the art on TIMIT phoneme classification and WSJ speech recognition.

Record transparency

Publication details

OpenAlex
W2996383576
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.