Alexei Baevski
7 papers in the PaperMetrix corpus
Papers by this author
-
Facebook FAIR’s WMT19 News Translation Task Submission
2019
This paper describes Facebook FAIR's submission to the WMT19 shared news translation task. We participate in four language directions, English German and English Russian in both directions. Following our submission from last year, our baseline …
-
Multilingual Speech Translation with Efficient Finetuning of Pretrained Models
2020 · arXiv (Cornell University)
We present a simple yet effective approach to build multilingual speech-to-text (ST) translation by efficient transfer learning from pretrained speech encoder and text decoder. Our key finding is that a minimalistic LNA (LayerNorm and Attention) …
-
Pay Less Attention with Lightweight and Dynamic Convolutions
2019 · arXiv (Cornell University)
Self-attention is a useful mechanism to build generative models for language and images. It determines the importance of context elements by comparing each element to the current time step. In this paper, we show that …
-
fairseq: A Fast, Extensible Toolkit for Sequence Modeling
2019
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, Michael Auli. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations). 2019.
-
vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
2019 · ArXiv.org
We propose vq-wav2vec to learn discrete representations of audio segments through a wav2vec-style self-supervised context prediction task. The algorithm uses either a gumbel softmax or online k-means clustering to quantize the dense representations. Discretization enables …
-
vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
2020 · arXiv (Cornell University)
We propose vq-wav2vec to learn discrete representations of audio segments through a wav2vec-style self-supervised context prediction task. The algorithm uses either a gumbel softmax or online k-means clustering to quantize the dense representations. Discretization enables …
-
XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale
2022 · Interspeech 2022
This paper presents XLS-R, a large-scale model for cross-lingual speech representation learning based on wav2vec 2.0.We train models with up to 2B parameters on nearly half a million hours of publicly available speech audio in …