ملف الباحث
Jiasheng Zheng
ورقتان في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models
2024 · arXiv (Cornell University)
In recent years, researchers have proposed numerous benchmarks to evaluate the impressive coding capabilities of large language models (LLMs). However, current benchmarks primarily assess the accuracy of LLM-generated code, while neglecting other critical dimensions that …
-
Multi-Facet Counterfactual Learning for Content Quality Evaluation
2024 · arXiv (Cornell University)
Evaluating the quality of documents is essential for filtering valuable content from the current massive amount of information. Conventional approaches typically rely on a single score as a supervision signal for training content quality evaluators, …