ملف الباحث

Wen-Ding Li

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

    2024 · arXiv (Cornell University)

    Large Language Models (LLMs) applied to code-related applications have emerged as a prominent field, attracting significant interest from both academia and industry. However, as new and improved LLMs are developed, existing evaluation benchmarks (e.g., HumanEval, …