conference-paper
وصول مفتوح
Reliable Confidence Intervals for Information Retrieval Evaluation Using Generative A.I.
Research footprint
At a glance
- الاستشهادات
- 2
- المراجع
- 44
- Comments
- 0
Paper overview
Abstract
The traditional evaluation of information retrieval (IR) systems is generally very costly as it requires manual relevance annotation from human experts. Recent advancements in generative artificial intelligence -specifically large language models (LLMs)- can generate relevance annotations at an enormous scale with relatively small computational costs. Potentially, this could alleviate the costs traditionally associated with IR evaluation and make it applicable to numerous low-resource applications. However, generated relevance annotations are not immune to (systematic) errors, and as a result, directly using them for evaluation produces unreliable results.
Record transparency
Publication details
- DOI
- 10.1145/3637528.3671883
- OpenAlex
- W4400374549
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.