Using ChatGPT to Score Essays and Short-Form Constructed Responses

Mark D. Shermis

arXiv:2408.09540·cs.CL·August 20, 2024

Using ChatGPT to Score Essays and Short-Form Constructed Responses

Mark D. Shermis

PDF

Open Access

TL;DR

This study evaluates ChatGPT's ability to score essays and responses, comparing its performance to human raters and machine models, highlighting its potential and current limitations for automated scoring.

Contribution

It provides an empirical assessment of ChatGPT's scoring accuracy across different models and datasets, identifying areas for improvement in fairness and reliability.

Findings

01

ChatGPT's gradient boost model achieved near-human QWK scores on some datasets.

02

Overall performance of ChatGPT was inconsistent and often below human scoring.

03

Further refinement is needed for ChatGPT to be reliable in high-stakes assessments.

Abstract

This study aimed to determine if ChatGPT's large language models could match the scoring accuracy of human and machine scores from the ASAP competition. The investigation focused on various prediction models, including linear regression, random forest, gradient boost, and boost. ChatGPT's performance was evaluated against human raters using quadratic weighted kappa (QWK) metrics. Results indicated that while ChatGPT's gradient boost model achieved QWKs close to human raters for some data sets, its overall performance was inconsistent and often lower than human scores. The study highlighted the need for further refinement, particularly in handling biases and ensuring scoring fairness. Despite these challenges, ChatGPT demonstrated potential for scoring efficiency, especially with domain-specific fine-tuning. The study concludes that ChatGPT can complement human scoring but requires…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsArtificial Intelligence in Healthcare and Education · Online Learning and Analytics