Towards Human-Level Evaluation: Assessing the Potential of GPT-4 in Automated Evaluation and Feedback Generation on Japanese Essays
At a glance
- Citations
- 3
- References
- 21
- Comments
- 0
Abstract
In recent years, Automated Writing Evaluation (AWE) has been extensively researched within the field of AI in education. This paper explores generative AI, such as GPT-4, which has garnered significant attention for its ability to score essays and provide feedback to students. We designed prompts for GPT-4 to assign scores and rationales based on a given rubric and to generate feedback beneficial for students' development. We compared the evaluations produced by GPT-4 with those made by human evaluators. The results demonstrate GPT-4's potential to assist in generating evaluations at a human level. In addition, we analysed the consistency of the scoring and the quality of the rationales and feedback generated by GPT-4. In this paper, we will share our analysis and also describe the points that need to be improved for implementation in practice.
Publication details
- DOI
- 10.1109/iiai-aai63651.2024.00039
- OpenAlex
- W4403421016
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.