ملف الباحث
Ruotong Pan
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models
2024 · arXiv (Cornell University)
In recent years, researchers have proposed numerous benchmarks to evaluate the impressive coding capabilities of large language models (LLMs). However, current benchmarks primarily assess the accuracy of LLM-generated code, while neglecting other critical dimensions that …