Researcher profile

Yinfei Yang

13 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. A Simple and Effective Method To Eliminate the Self Language Bias in Multilingual Representations

    2021 · Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing

    Language agnostic and semantic-language information isolation is an emerging research direction for multilingual representations models. We explore this problem from a novel angle of geometric algebra and semantic space. A simple but highly effective method …

  2. Language-agnostic BERT Sentence Embedding

    2020 · arXiv (Cornell University)

    While BERT is an effective method for learning monolingual sentence embeddings for semantic similarity and embedding based transfer learning (Reimers and Gurevych, 2019), BERT based cross-lingual sentence embeddings have yet to be explored. We systematically …

  3. Multimodal Autoregressive Pre-training of Large Vision Encoders

    2024 · arXiv (Cornell University)

    We introduce a novel method for pre-training of large-scale vision encoders. Building on recent advancements in autoregressive pre-training of vision models, we extend this framework to a multimodal setting, i.e., images and text. In this …

  4. Universal Sentence Encoder

    2018 · arXiv (Cornell University)

    We present models for encoding sentences into embedding vectors that specifically target transfer learning to other NLP tasks. The models are efficient and result in accurate performance on diverse transfer tasks. Two variants of the …

  5. Effective Parallel Corpus Mining using Bilingual Sentence Embeddings

    2018

    Mandy Guo, Qinlan Shen, Yinfei Yang, Heming Ge, Daniel Cer, Gustavo Hernandez Abrego, Keith Stevens, Noah Constant, Yun-Hsuan Sung, Brian Strope, Ray Kurzweil. Proceedings of the Third Conference on Machine Translation: Research Papers. 2018.

  6. Universal Sentence Encoder for English

    2018

    Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Brian Strope, Ray Kurzweil. Proceedings of the 2018 Conference on Empirical Methods in Natural …

  7. Learning Semantic Textual Similarity from Conversations

    2018

    Yinfei Yang, Steve Yuan, Daniel Cer, Sheng-yi Kong, Noah Constant, Petr Pilar, Heming Ge, Yun-Hsuan Sung, Brian Strope, Ray Kurzweil. Proceedings of the Third Workshop on Representation Learning for NLP. 2018.

  8. PAWS-X: A Cross-lingual Adversarial Dataset for Paraphrase Identification

    2019

    Yinfei Yang, Yuan Zhang, Chris Tar, Jason Baldridge. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.

  9. Multilingual Universal Sentence Encoder for Semantic Retrieval

    2020

    Yinfei Yang, Daniel Cer, Amin Ahmad, Mandy Guo, Jax Law, Noah Constant, Gustavo Hernandez Abrego, Steve Yuan, Chris Tar, Yun-hsuan Sung, Brian Strope, Ray Kurzweil. Proceedings of the 58th Annual Meeting of the Association for …

  10. Language-agnostic BERT Sentence Embedding

    2022 · Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

    While BERT is an effective method for learning monolingual sentence embeddings for semantic similarity and embedding based transfer learning (Reimers and Gurevych, 2019), BERT based cross-lingual sentence embeddings have yet to be explored. We systematically …

  11. Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models

    2022 · Findings of the Association for Computational Linguistics: ACL 2022

    We provide the first exploration of sentence embeddings from text-to-text transformers (T5) including the effects of scaling up sentence encoders to 11B parameters. Sentence embeddings are broadly useful for language processing tasks. While T5 achieves …

  12. LongT5: Efficient Text-To-Text Transformer for Long Sequences

    2022 · Findings of the Association for Computational Linguistics: NAACL 2022

    Recent work has shown that either (1) increasing the input length or (2) increasing model size can improve the performance of Transformer-based neural models. In this paper, we present LongT5, a new model that explores …

  13. Large Dual Encoders Are Generalizable Retrievers

    2022

    Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith Hall, Ming-Wei Chang, Yinfei Yang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. …