ملف الباحث

Zhicheng Dou

9 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Pchatbot: A Large-Scale Dataset for Personalized Chatbot

    2021

    atural language dialogue systems raise great attention recently. As many dialogue models are data-driven, high-quality datasets are essential to these systems. In this paper, we introduce Pchatbot, a large-scale dialogue dataset that contains two subsets …

  2. Less is More: Pre-train a Strong Text Encoder for Dense Retrieval Using a Weak Decoder

    2021 · arXiv (Cornell University)

    Dense retrieval requires high-quality text sequence embeddings to support effective search in the representation space. Autoencoder-based language models are appealing in dense retrieval as they train the encoder to output high-quality embedding that can reconstruct …

  3. Learning to Select Historical News Articles for Interaction based Neural News Recommendation

    2021 · arXiv (Cornell University)

    The key to personalized news recommendation is to match the user's interests with the candidate news precisely and efficiently. Most existing approaches embed user interests into a representation vector then recommend by comparing it with …

  4. Large Language Models Know Your Contextual Search Intent: A Prompting Framework for Conversational Search

    2023 · arXiv (Cornell University)

    Precisely understanding users' contextual search intent has been an important challenge for conversational search. As conversational search sessions are much more diverse and long-tailed, existing methods trained on limited data still show unsatisfactory effectiveness and …

  5. CL4DIV: A Contrastive Learning Framework for Search Result Diversification

    2024

    Search result diversification aims to provide a diversified document ranking list so as to cover as many intents as possible and satisfy the various information needs of different users. Existing approaches usually represented documents by …

  6. Generating Multi-turn Clarification for Web Information Seeking

    2024

    Asking multi-turn clarifying questions has been applied in various conversational search systems to help recommend people, commodities, and images to users. However, its importance is still not emphasized in the Web search. In this paper, …

  7. YuLan: An Open-source Large Language Model

    2024 · arXiv (Cornell University)

    Large language models (LLMs) have become the foundation of many applications, leveraging their extensive capabilities in processing and understanding natural language. While many open-source LLMs have been released with technical reports, the lack of training …

  8. Progressive Multimodal Reasoning via Active Retrieval

    2024 · arXiv (Cornell University)

    Multi-step multimodal reasoning tasks pose significant challenges for multimodal large language models (MLLMs), and finding effective ways to enhance their performance in such scenarios remains an unresolved issue. In this paper, we propose AR-MCTS, a …

  9. Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation

    2025

    Retrieval-augmented generation (RAG) has effectively mitigated the hallucination problem of large language models (LLMs). However, the difficulty of aligning the retriever with the LLMs' diverse knowledge preferences inevitably poses a challenge in developing a reliable …