preprint وصول مفتوح

Learning to Effectively Select Topics For Information Retrieval Test Collections.

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

الاستشهادات
1
المراجع
0
Comments
0
Paper overview

Abstract

While test collections are commonly employed to evaluate the effectiveness of information retrieval (IR) systems, constructing these collections has become increasingly expensive in recent years as collection sizes have grown ever-larger. To address this, we propose a new learning-to-rank topic selection method which reduces the number of search topics needed for reliable evaluation of IR systems. As part of this work, we also revisit the deep vs. shallow judging debate: whether it is better to collect many relevance judgments for a few topics or a few judgments for many topics. We consider a number of factors impacting this trade-off: how topics are selected, topic familiarity to judges, and how topic generation cost may impact both budget utilization and the resultant quality of judgments. Experiments on NIST TREC Robust 2003 and Robust 2004 test collections show not only our method's ability to reliably evaluate IR systems using fewer topics, but also that when topics are intelligently selected, deep judging is often more cost-effective than shallow judging in achieving the same level of evaluation reliability. Topic familiarity and construction costs are also seen to notably impact the evaluation cost vs. reliability tradeoff and provide further evidence supporting deep judging in practice.

Record transparency

Publication details

OpenAlex
W2580964469
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.