Introducing HALC: A general pipeline for finding optimal prompting strategies for automated coding with LLMs in the computational social sciences

Andreas Reich; Claudia Thoms; Tobias Schrimpf

arXiv:2507.21831·cs.CL·July 30, 2025

Introducing HALC: A general pipeline for finding optimal prompting strategies for automated coding with LLMs in the computational social sciences

Andreas Reich, Claudia Thoms, Tobias Schrimpf

PDF

TL;DR

This paper introduces HALC, a systematic pipeline for optimizing prompts in LLMs to improve automated coding accuracy in social science research, validated through extensive testing and expert comparison.

Contribution

HALC provides a general, reliable method for constructing optimal prompts across various tasks and models, reducing trial-and-error in LLM-based coding.

Findings

01

Prompts achieved high reliability with alpha > 0.7

02

Effective prompts identified for single and multiple variable coding

03

Insights into factors influencing prompt effectiveness

Abstract

LLMs are seeing widespread use for task automation, including automated coding in the social sciences. However, even though researchers have proposed different prompting strategies, their effectiveness varies across LLMs and tasks. Often trial and error practices are still widespread. We propose HALC $-$ a general pipeline that allows for the systematic and reliable construction of optimal prompts for any given coding task and model, permitting the integration of any prompting strategy deemed relevant. To investigate LLM coding and validate our pipeline, we sent a total of 1,512 individual prompts to our local LLMs in over two million requests. We test prompting strategies and LLM task performance based on few expert codings (ground truth). When compared to these expert codings, we find prompts that code reliably for single variables ( $α$ climate = .76; $α$ movement = .78) and…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.