Comparison of ChatGPT-4o and expert evaluation in endodontic education: a cross-sectional pilot study

Suha Alpay; Yasemen Darafarin; Burcu Dagdelen; Sana Mahroos Mkhailef Al-Shammari; Isil Kaya Buyukbayram

PMC · DOI:10.1186/s12909-025-08358-2·December 1, 2025

Comparison of ChatGPT-4o and expert evaluation in endodontic education: a cross-sectional pilot study

Suha Alpay, Yasemen Darafarin, Burcu Dagdelen, Sana Mahroos Mkhailef Al-Shammari, Isil Kaya Buyukbayram

PDF

Open Access

TL;DR

This study compared AI (ChatGPT-4o) and expert evaluations in assessing dental students' endodontic skills and found that while AI feedback was moderately useful, expert evaluations were rated higher.

Contribution

The study introduces a novel comparison of AI and expert evaluations in endodontic education and explores student perceptions of AI feedback.

Findings

01

AI and expert evaluations showed limited agreement, with ICC ranging from 0.36 to 0.45.

02

Students rated expert feedback higher than AI feedback for educational value and reliability.

03

Most students preferred a combination of AI and expert feedback for optimal learning.

Abstract

Artificial intelligence (AI) has the potential to enhance objectivity and scalability in educational assessment, yet its role in evaluating technical dental skills remains unclear. This study aimed to compare ChatGPT-4o based assessments with expert evaluations in undergraduate endodontic training and to explore student perceptions of AI-assisted feedback. This cross-sectional pilot study was conducted during the 2024–2025 academic year with 32 dental students from a faculty of dentistry, who completed root canal treatments. Postoperative radiographs were evaluated by 10 years experienced endodontist and ChatGPT-4o was used to evaluate performance based on five standardized criteria: canal centering, homogeneity, procedural errors, apical shaping, and overall taper, each rated on a 5-point Likert scale. Inter-rater reliability was assessed via intraclass correlation coefficients (ICC),…

Linked entities

Genes, proteins, chemicals, diseases, species, mutations and cell lines named across the full text — each resolved to its canonical identifier and authoritative record.

Genes1

SHROOM4

Proteins1

Species1

Homo sapiens(human · species)

Chemicals1

GPT-4 V

Diseases5

LLM AI hallucinations supernumerary teeth XAI

Figures6

Click any figure to enlarge with its caption.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsArtificial Intelligence in Healthcare and Education · Clinical Reasoning and Diagnostic Skills · Explainable Artificial Intelligence (XAI)