article Open access

Accuracy and reliability of large language models in assessing learning outcomes achievement across cognitive domains

  • AJP Advances in Physiology Education
  • American Physical Society
Research footprint

At a glance

Citations
13
References
22
Comments
0
Paper overview

Abstract

The advent of large language models (LLMs) such as ChatGPT and Gemini has offered new learning and assessment opportunities to integrate artificial intelligence (AI) with education. This study evaluated the accuracy of LLMs in assessing an assignment from a course on sports physiology. Concordance and correlation between human graders and LLMs were mostly moderate to poor. The findings suggest AI's potential to complement human expertise in educational assessment alongside the need for adaptive learning by educators.

Record transparency

Publication details

DOI
10.1152/advan.00137.2024
OpenAlex
W4404174416
Document type
article
Language
EN
Source
AJP Advances in Physiology Education
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.