conference-paper Open access

Towards Human-Level Evaluation: Assessing the Potential of GPT-4 in Automated Evaluation and Feedback Generation on Japanese Essays

Research footprint

At a glance

Citations
3
References
21
Comments
0
Paper overview

Abstract

In recent years, Automated Writing Evaluation (AWE) has been extensively researched within the field of AI in education. This paper explores generative AI, such as GPT-4, which has garnered significant attention for its ability to score essays and provide feedback to students. We designed prompts for GPT-4 to assign scores and rationales based on a given rubric and to generate feedback beneficial for students' development. We compared the evaluations produced by GPT-4 with those made by human evaluators. The results demonstrate GPT-4's potential to assist in generating evaluations at a human level. In addition, we analysed the consistency of the scoring and the quality of the rationales and feedback generated by GPT-4. In this paper, we will share our analysis and also describe the points that need to be improved for implementation in practice.

Record transparency

Publication details

DOI
10.1109/iiai-aai63651.2024.00039
OpenAlex
W4403421016
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.