Zhicheng Dou
9 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Pchatbot: A Large-Scale Dataset for Personalized Chatbot
2021
atural language dialogue systems raise great attention recently. As many dialogue models are data-driven, high-quality datasets are essential to these systems. In this paper, we introduce Pchatbot, a large-scale dialogue dataset that contains two subsets …
-
Less is More: Pre-train a Strong Text Encoder for Dense Retrieval Using a Weak Decoder
2021 · arXiv (Cornell University)
Dense retrieval requires high-quality text sequence embeddings to support effective search in the representation space. Autoencoder-based language models are appealing in dense retrieval as they train the encoder to output high-quality embedding that can reconstruct …
-
Learning to Select Historical News Articles for Interaction based Neural News Recommendation
2021 · arXiv (Cornell University)
The key to personalized news recommendation is to match the user's interests with the candidate news precisely and efficiently. Most existing approaches embed user interests into a representation vector then recommend by comparing it with …
-
Large Language Models Know Your Contextual Search Intent: A Prompting Framework for Conversational Search
2023 · arXiv (Cornell University)
Precisely understanding users' contextual search intent has been an important challenge for conversational search. As conversational search sessions are much more diverse and long-tailed, existing methods trained on limited data still show unsatisfactory effectiveness and …
-
CL4DIV: A Contrastive Learning Framework for Search Result Diversification
2024
Search result diversification aims to provide a diversified document ranking list so as to cover as many intents as possible and satisfy the various information needs of different users. Existing approaches usually represented documents by …
-
Generating Multi-turn Clarification for Web Information Seeking
2024
Asking multi-turn clarifying questions has been applied in various conversational search systems to help recommend people, commodities, and images to users. However, its importance is still not emphasized in the Web search. In this paper, …
-
YuLan: An Open-source Large Language Model
2024 · arXiv (Cornell University)
Large language models (LLMs) have become the foundation of many applications, leveraging their extensive capabilities in processing and understanding natural language. While many open-source LLMs have been released with technical reports, the lack of training …
-
Progressive Multimodal Reasoning via Active Retrieval
2024 · arXiv (Cornell University)
Multi-step multimodal reasoning tasks pose significant challenges for multimodal large language models (MLLMs), and finding effective ways to enhance their performance in such scenarios remains an unresolved issue. In this paper, we propose AR-MCTS, a …
-
Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation
2025
Retrieval-augmented generation (RAG) has effectively mitigated the hallucination problem of large language models (LLMs). However, the difficulty of aligning the retriever with the LLMs' diverse knowledge preferences inevitably poses a challenge in developing a reliable …