RED-CT: A Systems Design Methodology for Using LLM-labeled Data to Train   and Deploy Edge Classifiers for Computational Social Science

David Farr; Nico Manzonelli; Iain Cruickshank; and Jevin West

arXiv:2408.08217·cs.LG·November 5, 2024·3 cites

RED-CT: A Systems Design Methodology for Using LLM-labeled Data to Train and Deploy Edge Classifiers for Computational Social Science

David Farr, Nico Manzonelli, Iain Cruickshank, and Jevin West

PDF

Open Access

TL;DR

This paper presents RED-CT, a systems design methodology that leverages large language models as imperfect data annotators, with intervention measures to improve classification performance for deploying edge classifiers in social science applications.

Contribution

It introduces a novel systems design approach that effectively integrates LLMs into supervised learning workflows, outperforming LLM-generated labels in most tests.

Findings

01

Method outperforms LLM-generated labels in 7 of 8 tests.

02

System intervention measures improve classification accuracy.

03

Applicable to deploying edge classifiers in social science contexts.

Abstract

Large language models (LLMs) have enhanced our ability to rapidly analyze and classify unstructured natural language data. However, concerns regarding cost, network limitations, and security constraints have posed challenges for their integration into work processes. In this study, we adopt a systems design approach to employing LLMs as imperfect data annotators for downstream supervised learning tasks, introducing novel system intervention measures aimed at improving classification performance. Our methodology outperforms LLM-generated labels in seven of eight tests, demonstrating an effective strategy for incorporating LLMs into the design and deployment of specialized, supervised learning models present in many industry use cases.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsComputational and Text Analysis Methods