Comparing large language models for supervised analysis of students' lab   notes

Rebeckah K. Fussell; Megan Flynn; Anil Damle; Michael F.J. Fox; N. G.; Holmes

arXiv:2412.10610·physics.ed-ph·February 25, 2025

Comparing large language models for supervised analysis of students' lab notes

Rebeckah K. Fussell, Megan Flynn, Anil Damle, Michael F.J. Fox, N. G., Holmes

PDF

Open Access

TL;DR

This study compares various large language models and traditional methods for analyzing students' lab notes, focusing on their performance, resource use, and research implications in physics education.

Contribution

It provides a comparative analysis of fine-tuned and few-shot LLMs versus traditional methods for classifying student lab notes in physics education research.

Findings

01

Higher-resource models often perform better but not always.

02

All models show similar research trend estimations.

03

Absolute measurement values vary beyond uncertainties.

Abstract

Recent advancements in large language models (LLMs) hold significant promise in improving physics education research that uses machine learning. In this study, we compare the application of various models to perform large-scale analysis of written text grounded in a physics education research classification problem: identifying skills in students' typed lab notes through sentence-level labeling. Specifically, we use training data to fine-tune two different LLMs, BERT and LLaMA, and compare the performance of these models to both a traditional bag of words approach and a few-shot LLM (without fine-tuning).} We evaluate the models based on their resource use, performance metrics, and research outcomes when identifying skills in lab notes. We find that higher-resource models often, but not necessarily, perform better than lower-resource models. We also find that all models estimate similar…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsOnline Learning and Analytics