Pith. sign in

REVIEW 1 cited by

Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.17633 v1 pith:COLZ2L7Q submitted 2024-06-25 cs.CL cs.LG

classification cs.CLcs.LG
keywords labelssupervisedclassificationclassifiersdatafine-tunedmodelstext
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Computational social science (CSS) practitioners often rely on human-labeled data to fine-tune supervised text classifiers. We assess the potential for researchers to augment or replace human-generated training data with surrogate training labels from generative large language models (LLMs). We introduce a recommended workflow and test this LLM application by replicating 14 classification tasks and measuring performance. We employ a novel corpus of English-language text classification data sets from recent CSS articles in high-impact journals. Because these data sets are stored in password-protected archives, our analyses are less prone to issues of contamination. For each task, we compare supervised classifiers fine-tuned using GPT-4 labels against classifiers fine-tuned with human annotations and against labels from GPT-4 and Mistral-7B with few-shot in-context learning. Our findings indicate that supervised classification models fine-tuned on LLM-generated labels perform comparably to models fine-tuned with labels from human annotators. Fine-tuning models using LLM-generated labels can be a fast, efficient and cost-effective method of building supervised text classifiers.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Poodle: Seamlessly Scaling Down Large Language Models with Just-in-Time Model Replacement

    cs.DB 2025-12 unverdicted novelty 4.0 of 10

    Poodle shows that LLMs can be automatically replaced with cheaper models for recurring tasks to save significant cost and energy without extra user effort.

Pith tools