Pith. sign in

REVIEW 2 cited by

The Promises and Pitfalls of LLM Annotations in Dataset Labeling: a Case Study on Media Bias Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.11081 v2 pith:DCJABTA4 submitted 2024-11-17 cs.CL

classification cs.CL
keywords biasmediadatasetllmsdatadetectionannotatingannotation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

High annotation costs from hiring or crowdsourcing complicate the creation of large, high-quality datasets needed for training reliable text classifiers. Recent research suggests using Large Language Models (LLMs) to automate the annotation process, reducing these costs while maintaining data quality. LLMs have shown promising results in annotating downstream tasks like hate speech detection and political framing. Building on the success in these areas, this study investigates whether LLMs are viable for annotating the complex task of media bias detection and whether a downstream media bias classifier can be trained on such data. We create annolexical, the first large-scale dataset for media bias classification with over 48000 synthetically annotated examples. Our classifier, fine-tuned on this dataset, surpasses all of the annotator LLMs by 5-9 percent in Matthews Correlation Coefficient (MCC) and performs close to or outperforms the model trained on human-labeled data when evaluated on two media bias benchmark datasets (BABE and BASIL). This study demonstrates how our approach significantly reduces the cost of dataset creation in the media bias domain and, by extension, the development of classifiers, while our subsequent behavioral stress-testing reveals some of its current limitations and trade-offs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Prosocial Behavior Detection in Player Game Chat: From Aligning Human-AI Definitions to Efficient Annotation at Scale

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A hybrid system using a calibrated SVM plus selective GPT-4o fallback detects prosocial game chat at roughly 0.90 precision while cutting LLM inference cost by about 70%.

  2. Are Large Language Models the future crowd workers of Linguistics?

    cs.CL 2025-02 conditional novelty 5.0 of 10

    In two replicated linguistics experiments, GPT-4o-mini's zero-shot responses matched or beat published human performance, but the study lacks statistical validation and relies on only two tasks.

Pith tools