Juan Pino
8 papers in the PaperMetrix corpus
Papers by this author
-
Multilingual Speech Translation with Efficient Finetuning of Pretrained Models
2020 · arXiv (Cornell University)
We present a simple yet effective approach to build multilingual speech-to-text (ST) translation by efficient transfer learning from pretrained speech encoder and text decoder. Our key finding is that a minimalistic LNA (LayerNorm and Attention) …
-
A General Multi-Task Learning Framework to Leverage Text Data for Speech to Text Tasks
2021
Attention-based sequence-to-sequence modeling provides a powerful and elegant solution for applications that need to map one sequence to a different sequence. Its success heavily relies on the availability of large amounts of training data. This …
-
Pre-training for Speech Translation: CTC Meets Optimal Transport (2)
2023 · Zenodo (CERN European Organization for Nuclear Research)
Pre-trained models for the paper: Pre-training for Speech Translation: CTC Meets Optimal Transport. - MT models for MuST-C and CoVoST-2 - ASR and ST models for CoVoST (one-to-many and many-to-one)
-
Enhancing Speech-To-Speech Translation with Multiple TTS Targets
2023
It has been known that direct speech-to-speech translation (S2ST) models usually suffer from the data scarcity issue because of the limited existing parallel materials for both source and target speech. Therefore to train a direct …
-
Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text Tasks
2023 · arXiv (Cornell University)
Transducer and Attention based Encoder-Decoder (AED) are two widely used frameworks for speech-to-text tasks. They are designed for different purposes and each has its own benefits and drawbacks for speech-to-text tasks. In order to leverage …
-
The FLORES Evaluation Datasets for Low-Resource Machine Translation: Nepali–English and Sinhala–English
2019
Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, Marc’Aurelio Ranzato. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on …
-
On Evaluation of Adversarial Perturbations for Sequence-to-Sequence Models
2019
Paul Michel, Xian Li, Graham Neubig, Juan Pino. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 2019.
-
XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale
2022 · Interspeech 2022
This paper presents XLS-R, a large-scale model for cross-lingual speech representation learning based on wav2vec 2.0.We train models with up to 2B parameters on nearly half a million hours of publicly available speech audio in …