ملف الباحث
Tzu-Heng Huang
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
2025 · arXiv (Cornell University)
Large language models (LLMs) are widely used to evaluate the quality of LLM generations and responses, but this leads to significant challenges: high API costs, uncertain reliability, inflexible pipelines, and inherent biases. To address these, …