ملف الباحث
Yuekai Sun
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
tinyBenchmarks: evaluating LLMs with fewer examples
2024 · arXiv (Cornell University)
The versatility of large language models (LLMs) led to the creation of diverse benchmarks that thoroughly test a variety of language models' abilities. These benchmarks consist of tens of thousands of examples making evaluation of …