Researcher profile

Yi Tay

14 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Simple and Effective Curriculum Pointer-Generator Networks for Reading Comprehension over Long Narratives

    2019 · arXiv (Cornell University)

    This paper tackles the problem of reading comprehension over long narratives where documents easily span over thousands of tokens. We propose a curriculum learning (CL) based Pointer-Generator framework for reading/sampling over large documents, enabling diverse …

  2. Multi-Task Neural Network for Non-discrete Attribute Prediction in Knowledge Graphs

    2017

    Many popular knowledge graphs such as Freebase, YAGO or DBPedia maintain a list of non-discrete attributes for each entity. Intuitively, these attributes such as height, price or population count are able to richly characterize entities …

  3. HyperML

    2020

    This paper investigates the notion of learning user and item representations in non-Euclidean space. Specifically, we study the connection between metric learning in hyperbolic space and collaborative filtering by exploring Mobius gyrovector spaces where the …

  4. Knowledge Router: Learning Disentangled Representations for Knowledge Graphs

    2021

    The design of expressive representations of entities and relations in a knowledge graph is an important endeavor. While many of the existing approaches have primarily focused on learning from relational patterns and structural information, the …

  5. Emergent Abilities of Large Language Models

    2022 · arXiv (Cornell University)

    Scaling up language models has been shown to predictably improve performance and sample efficiency on a wide range of downstream tasks. This paper instead discusses an unpredictable phenomenon that we refer to as emergent abilities …

  6. Confident Adaptive Language Modeling

    2022 · arXiv (Cornell University)

    Recent advances in Transformer-based large language models (LLMs) have led to significant performance improvements across many tasks. These gains come with a drastic increase in the models' size, potentially leading to slow and costly use …

  7. Symbol tuning improves in-context learning in language models

    2023 · arXiv (Cornell University)

    We present symbol tuning - finetuning language models on in-context input-label pairs where natural language labels (e.g., "positive/negative sentiment") are replaced with arbitrary symbols (e.g., "foo/bar"). Symbol tuning leverages the intuition that when a model …

  8. Learning to Rank Question Answer Pairs with Holographic Dual LSTM Architecture

    2017

    We describe a new deep learning architecture for learning to rank question answer pairs. Our approach extends the long short-term memory (LSTM) network with holographic composition to model the relationship between question and answer representations. …

  9. Latent Relational Metric Learning via Memory-based Attention for Collaborative Ranking

    2018

    This paper proposes a new neural architecture for collaborative ranking with implicit feedback. Our model, LRML (Latent Relational Metric Learning) is a novel metric learning approach for recommendation. More specifically, instead of simple push-pull mechanisms …

  10. Multi-Pointer Co-Attention Networks for Recommendation

    2018

    Many recent state-of-the-art recommender systems such as D-ATT, TransNet and DeepCoNN exploit reviews for representation learning. This paper proposes a new neural architecture for recommendation with reviews. Our model operates on a multi-hierarchical paradigm and …

  11. PaLM: Scaling Language Modeling with Pathways

    2022 · arXiv (Cornell University)

    Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to …

  12. Deep Learning Based Recommender System

    2019 · ACM Computing Surveys

    With the growing volume of online information, recommender systems have been an effective strategy to overcome information overload. The utility of recommender systems cannot be overstated, given their widespread adoption in many web applications, along …

  13. UL2: Unifying Language Learning Paradigms

    2022 · arXiv (Cornell University)

    Existing pre-trained models are generally geared towards a particular class of problems. To date, there seems to be still no consensus on what the right architecture and pre-training setup should be. This paper presents a …

  14. Scaling Instruction-Finetuned Language Models

    2022 · arXiv (Cornell University)

    Finetuning language models on a collection of datasets phrased as instructions has been shown to improve model performance and generalization to unseen tasks. In this paper we explore instruction finetuning with a particular focus on …