Published January 1, 2025 | Version v1
Journal article Open

Assessing AI in Educational Evaluation: A Comprehensive Analysis of ChatGPT's Performance on PISA Reading Skills

  • 1. Gaziantep Univ, Fac Educ, Dept Educ Sci, Gaziantep, Turkiye
  • 2. Sakarya Univ, Fac Educ, Dept Educ Sci, Sakarya, Turkiye

Description

This study explores ChatGPT's ability to solve reading skill questions from the PISA exams, an internationally recognized assessment designed to evaluate 15-year-old students' reading, mathematics, and science competencies. 90 questions from the 2000 and 2009 PISA reading assessments were administered to ChatGPT twice within one week. Questions were analyzed based on their text type (continuous, non-continuous, or multiple), question format (open-ended, multiple-choice, or short answer), and difficulty level defined by the PISA framework. The accuracy of ChatGPT's responses was evaluated using the official answer keys provided by the OECD. Results indicate that ChatGPT achieved a 91% accuracy rate in both rounds. However, there were discrepancies in the questions answered incorrectly between the two rounds. ChatGPT performed better on continuous texts than non-continuous texts and achieved its highest accuracy rates on questions of lower difficulty. The analysis also highlights inconsistencies in ChatGPT's responses, particularly in handling non-continuous texts and higher-level questions based on Bloom's taxonomy. While ChatGPT has limitations, such as occasional incorrect or inconsistent answers, it remains a valuable tool with the potential for educational integration, particularly in automating assessment processes. The study underscores the importance of understanding ChatGPT's strengths and limitations to enhance its application in educational settings.

Files

bib-407b9703-98e9-4764-9039-69301445ca97.txt

Files (200 Bytes)

Name Size Download all
md5:ffa02c9875ee00f1366406ac49e8689a
200 Bytes Preview Download