conference-paper وصول مفتوح

Reliable Confidence Intervals for Information Retrieval Evaluation Using Generative A.I.

Research footprint

At a glance

الاستشهادات
2
المراجع
44
Comments
0
Paper overview

Abstract

The traditional evaluation of information retrieval (IR) systems is generally very costly as it requires manual relevance annotation from human experts. Recent advancements in generative artificial intelligence -specifically large language models (LLMs)- can generate relevance annotations at an enormous scale with relatively small computational costs. Potentially, this could alleviate the costs traditionally associated with IR evaluation and make it applicable to numerous low-resource applications. However, generated relevance annotations are not immune to (systematic) errors, and as a result, directly using them for evaluation produces unreliable results.

Record transparency

Publication details

DOI
10.1145/3637528.3671883
OpenAlex
W4400374549
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.