Exploring Large Language Model-Aided Document Expansion for Sparse Retrieval
At a glance
- Citations
- 0
- References
- 11
- Comments
- 0
Abstract
Document expansion is a technique designed to mitigate the term mismatch problem by enriching documents with related terms or queries, thereby improving retrieval performance. Notable approaches like doc2query and docT5query utilize sequence-to-sequence transformers or T5 models to generate queries from documents. However, these methods rely on a substantial number of relevant query-document pairs for training, which can be resource-intensive to obtain. Recently, generative instruction-tuned large language models (LLMs) have demonstrated remarkable capabilities across various language tasks, making them a promising alternative for document expansion without requiring additional training. To this end, this paper explores the potential of LLMs for document expansion, evaluating whether their advanced capabilities can enhance retrieval by generating more accurate and contextually rele- vant queries for documents. Experiments conducted on the MS MARCO and TREC DL datasets using Mistral and Vicuna models reveal that zero-shot document expansion using LLMs achieves retrieval performance comparable to in-domain trained T5 models. Meanwhile, LLMs could generate more diverse queries so that the document context may be shifted when adding generated queries too much. Furthermore, LLMs can generate more diverse queries, but excessive inclusion of these queries may shift the original document context, highlighting a trade-off in using LLMgenerated queries for document expansion. These findings underscore the potential of LLMs as an effective and resourceefficient solution for document expansion in information retrieval tasks.
Publication details
- DOI
- 10.1109/ainit65432.2025.11035404
- OpenAlex
- W4411552748
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.