Noah Constant
12 papers in the PaperMetrix corpus
Papers by this author
-
nmT5 -- Is parallel data still relevant for pre-training massively multilingual language models?
2021 · arXiv (Cornell University)
Recently, mT5 - a massively multilingual version of T5 - leveraged a unified text-to-text format to attain state-of-the-art results on a wide variety of multilingual NLP tasks. In this paper, we investigate the impact of …
-
Universal Sentence Encoder
2018 · arXiv (Cornell University)
We present models for encoding sentences into embedding vectors that specifically target transfer learning to other NLP tasks. The models are efficient and result in accurate performance on diverse transfer tasks. Two variants of the …
-
Effective Parallel Corpus Mining using Bilingual Sentence Embeddings
2018
Mandy Guo, Qinlan Shen, Yinfei Yang, Heming Ge, Daniel Cer, Gustavo Hernandez Abrego, Keith Stevens, Noah Constant, Yun-Hsuan Sung, Brian Strope, Ray Kurzweil. Proceedings of the Third Conference on Machine Translation: Research Papers. 2018.
-
Universal Sentence Encoder for English
2018
Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Brian Strope, Ray Kurzweil. Proceedings of the 2018 Conference on Empirical Methods in Natural …
-
Character-Level Language Modeling with Deeper Self-Attention
2019
LSTMs and other RNN variants have shown strong performance on character-level language modeling. These models are typically trained using truncated backpropagation through time, and it is common to assume that their success stems from their …
-
Learning Semantic Textual Similarity from Conversations
2018
Yinfei Yang, Steve Yuan, Daniel Cer, Sheng-yi Kong, Noah Constant, Petr Pilar, Heming Ge, Yun-Hsuan Sung, Brian Strope, Ray Kurzweil. Proceedings of the Third Workshop on Representation Learning for NLP. 2018.
-
Multilingual Universal Sentence Encoder for Semantic Retrieval
2020
Yinfei Yang, Daniel Cer, Amin Ahmad, Mandy Guo, Jax Law, Noah Constant, Gustavo Hernandez Abrego, Steve Yuan, Chris Tar, Yun-hsuan Sung, Brian Strope, Ray Kurzweil. Proceedings of the 58th Annual Meeting of the Association for …
-
mT5: A massively multilingual pre-trained text-to-text transformer
2020 · arXiv (Cornell University)
The recent "Text-to-Text Transfer Transformer" (T5) leveraged a unified text-to-text format and scale to attain state-of-the-art results on a wide variety of English-language NLP tasks. In this paper, we introduce mT5, a multilingual variant of …
-
ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models
2022 · Transactions of the Association for Computational Linguistics
Abstract Most widely used pre-trained language models operate on sequences of tokens corresponding to word or subword units. By comparison, token-free models that operate directly on raw text (bytes or characters) have many benefits: They …
-
mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer
2021
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, Colin Raffel. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. …
-
Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models
2022 · Findings of the Association for Computational Linguistics: ACL 2022
We provide the first exploration of sentence embeddings from text-to-text transformers (T5) including the effects of scaling up sentence encoders to 11B parameters. Sentence embeddings are broadly useful for language processing tasks. While T5 achieves …
-
SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer
2022 · Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
There has been growing interest in parameterefficient methods to apply pre-trained language models to downstream tasks. Building on the PROMPTTUNING approach of Lester et al. ( SPOT first learns a prompt on one or more …