Cunxiang Wang
3 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
SemEval-2020 Task 4: Commonsense Validation and Explanation
2020
In this paper, we present SemEval-2020 Task 4, Commonsense Validation and Explanation (ComVE), which includes three subtasks, aiming to evaluate whether a system can distinguish a natural language statement that makes sense to humans from …
-
Evaluating Open-QA Evaluation
2023 · arXiv (Cornell University)
This study focuses on the evaluation of the Open Question Answering (Open-QA) task, which can directly estimate the factuality of large language models (LLMs). Current automatic evaluation methods have shown limitations, indicating that human evaluation …
-
A Survey on Evaluation of Large Language Models
2024 · ACM Transactions on Intelligent Systems and Technology
Large language models (LLMs) are gaining increasing popularity in both academia and industry, owing to their unprecedented performance in various applications. As LLMs continue to play a vital role in both research and daily use, …