ملف الباحث

Yuekai Sun

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. tinyBenchmarks: evaluating LLMs with fewer examples

    2024 · arXiv (Cornell University)

    The versatility of large language models (LLMs) led to the creation of diverse benchmarks that thoroughly test a variety of language models' abilities. These benchmarks consist of tens of thousands of examples making evaluation of …