ملف الباحث

Zhengzhao Ma

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models

    2024 · arXiv (Cornell University)

    In recent years, researchers have proposed numerous benchmarks to evaluate the impressive coding capabilities of large language models (LLMs). However, current benchmarks primarily assess the accuracy of LLM-generated code, while neglecting other critical dimensions that …