Researcher profile

Yutao Zhu

7 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Pchatbot: A Large-Scale Dataset for Personalized Chatbot

    2021

    atural language dialogue systems raise great attention recently. As many dialogue models are data-driven, high-quality datasets are essential to these systems. In this paper, we introduce Pchatbot, a large-scale dialogue dataset that contains two subsets …

  2. An Empirical Study of Uniform-Architecture Knowledge Distillation in Document Ranking

    2023 · arXiv (Cornell University)

    Although BERT-based ranking models have been commonly used in commercial search engines, they are usually time-consuming for online ranking tasks. Knowledge distillation, which aims at learning a smaller model with comparable performance to a larger …

  3. CL4DIV: A Contrastive Learning Framework for Search Result Diversification

    2024

    Search result diversification aims to provide a diversified document ranking list so as to cover as many intents as possible and satisfy the various information needs of different users. Existing approaches usually represented documents by …

  4. YuLan: An Open-source Large Language Model

    2024 · arXiv (Cornell University)

    Large language models (LLMs) have become the foundation of many applications, leveraging their extensive capabilities in processing and understanding natural language. While many open-source LLMs have been released with technical reports, the lack of training …

  5. Progressive Multimodal Reasoning via Active Retrieval

    2024 · arXiv (Cornell University)

    Multi-step multimodal reasoning tasks pose significant challenges for multimodal large language models (MLLMs), and finding effective ways to enhance their performance in such scenarios remains an unresolved issue. In this paper, we propose AR-MCTS, a …

  6. Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation

    2025

    Retrieval-augmented generation (RAG) has effectively mitigated the hallucination problem of large language models (LLMs). However, the difficulty of aligning the retriever with the LLMs' diverse knowledge preferences inevitably poses a challenge in developing a reliable …

  7. S3-Rec: Self-Supervised Learning for Sequential Recommendation with Mutual Information Maximization

    2020

    Recently, significant progress has been made in sequential recommendation with deep learning. Existing neural sequential recommendation models usually rely on the item prediction loss to learn model parameters or data representations. However, the model trained …