Yuchi Ma
3 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
CoderEval: A Benchmark of Pragmatic Code Generation with Generative Pre-trained Models
2024
Code generation models based on the pre-training and fine-tuning paradigm have been increasingly attempted by both academia and industry, resulting in well-known industrial models such as Codex, CodeGen, and PanGu-Coder. To evaluate the effectiveness of …
-
HumanEvo: An Evolution-Aware Benchmark for More Realistic Evaluation of Repository-Level Code Generation
2025
To evaluate the repository-level code generation capabilities of Large Language Models (LLMs) in complex real-world software development scenarios, many evaluation methods have been developed. These methods typically leverage contextual code from the latest version of …
-
Towards Mitigating API Hallucination in Code Generated by LLMs with Hierarchical Dependency Aware
2025 · arXiv (Cornell University)
Application Programming Interfaces (APIs) are crucial in modern software development. Large Language Models (LLMs) assist in automated code generation but often struggle with API hallucination, including invoking non-existent APIs and misusing existing ones in practical …