ملف الباحث
Sher Badshah
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
DAFE: LLM-Based Evaluation Through Dynamic Arbitration for Free-Form Question-Answering
2025 · Qeios
Evaluating Large Language Models (LLMs) free-form generated responses remains a challenge due to their diverse and open-ended nature. Traditional supervised signal-based automatic metrics fail to capture semantic equivalence or handle the variability of open-ended responses, …