Pith. sign in

REVIEW 15 cited by

Automated Annotation with Generative AI Requires Validation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.00176 v1 pith:WI6OAYJF submitted 2023-05-31 cs.CL cs.AI

classification cs.CLcs.AI
keywords annotationautomatedllmsperformancetextvalidateacrossgenerative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative large language models (LLMs) can be a powerful tool for augmenting text annotation procedures, but their performance varies across annotation tasks due to prompt quality, text data idiosyncrasies, and conceptual difficulty. Because these challenges will persist even as LLM technology improves, we argue that any automated annotation process using an LLM must validate the LLM's performance against labels generated by humans. To this end, we outline a workflow to harness the annotation potential of LLMs in a principled, efficient way. Using GPT-4, we validate this approach by replicating 27 annotation tasks across 11 datasets from recent social science articles in high-impact journals. We find that LLM performance for text annotation is promising but highly contingent on both the dataset and the type of annotation task, which reinforces the necessity to validate on a task-by-task basis. We make available easy-to-use software designed to implement our workflow and streamline the deployment of LLMs for automated annotation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Talking Politics with Artificial Intelligence

    econ.GN 2026-07 unverdicted novelty 7.0 of 10

    Large-scale analysis of AI conversations indicates they function primarily as practical intermediaries for political tasks rather than arenas for public expression, with increased expressiveness after major events.

  2. The Model as One Rater Among Several: Measuring Political Positions in Data-Sparse Regions with a Language-Model Panel

    cs.CY 2026-06 unverdicted novelty 7.0 of 10

    A panel of nine LLMs achieves Krippendorff's alpha of 0.86 for political position measurement in data-sparse regions, with added axis definitions improving agreement and disagreements revealing interpretive issues.

  3. LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification

    cs.CL 2026-04 conditional novelty 7.0 of 10

    Fine-tuned BERTimbau-LoRA achieves 87.6% accuracy and 0.87 macro-F1 on LegalBench-BR, outperforming commercial LLMs by 22-28 points and eliminating their systematic bias toward civil law on Brazilian legal classification.

  4. Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection

    cs.CL 2026-04 conditional novelty 7.0 of 10

    LLM annotation can replace human labels for hostility detection with comparable F1 at much lower cost, but active learning adds little value and error structures differ systematically.

  5. Guidelines for Empirical Studies in Software Engineering involving Large Language Models

    cs.SE 2025-08 accept novelty 7.0 of 10

    The paper delivers a taxonomy of seven LLM study types in software engineering along with eight guidelines that separate mandatory requirements from recommended practices to address reproducibility challenges.

  6. Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection

    cs.CL 2026-04 conditional novelty 6.5 of 10

    With a two-question interface, full-pool LLM annotation (GPT-5.2 or Qwen3.5) beats human-supervised classifiers for anti-immigrant hostility detection at ~1/10 cost; AL does not reliably beat random sampling.

  7. Auditing Differential Visibility of Political Content on TikTok

    cs.SI 2026-07 conditional novelty 6.0 of 10

    Account-level analysis finds no evidence of moderate-to-large reach suppression on TikTok for three political topics; an apparent pooled gap is a statistical artifact, while oppositional content earns more engagement ...

  8. Grounded Event Extraction from SEC 8-K Filings with a Fine-Grained Taxonomy

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Schema-constrained, quote-grounded LLM extraction plus a second-pass quality score yields 601k auditable 8-K event tags whose precision and market reactions both improve with the score.

  9. Correct codes for the wrong reasons? validating LLMs as measurement instruments for theoretical constructs

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    Grain calibration decomposes theoretical constructs into clause-level components, tests each with extractive evidence, and combines results through explicit theory-derived rules to validate LLM coding beyond agreement...

  10. What Your Posts Reveal: A Benchmark and Agentic Framework for User-Level Privacy Leakage on Social Media

    cs.CR 2026-06 unverdicted novelty 6.0 of 10

    Introduces SopriBench benchmark and Argus agentic framework for user-level multimodal privacy leakage inference, reporting 0.55 PES with 25% improvement over baselines.

  11. Guidelines for Empirical Studies in Software Engineering involving Large Language Models

    cs.SE 2025-08 accept novelty 6.0 of 10

    A group of 22 researchers proposes seven study types and eight guidelines for empirical software engineering studies involving LLMs to enhance reproducibility and replicability.

  12. A Durability and Cross-Language Transfer Benchmark for a Validated Teaching-Feedback Classification Protocol

    cs.CL 2026-07 accept novelty 5.0 of 10

    The teaching-feedback classification protocol remains durable across three representation generations and transfers to English sentiment, so model choice is a deployment decision.

  13. Talking Politics with Artificial Intelligence

    econ.GN 2026-07 unverdicted novelty 5.0 of 10

    Political content appears in 3.9% of AI conversations, mostly for information and drafting rather than opinions, with U.S. users showing increased stance-taking and affect after the 2024 election result call.

  14. How Much Does Persuasion Strategy Matter? LLM-Annotated Evidence from Charitable Donation Dialogues

    cs.CL 2026-03 unverdicted novelty 5.0 of 10

    LLM annotation of donation dialogues shows persuasion strategy categories explain only 1.5% of outcome variance, with guilt induction linked to 23 percentage point lower donation rates that replicate across models.

  15. LLM Predictive Scoring and Validation: Inferring Experience Ratings from Unstructured Text

    cs.CL 2026-04 unverdicted novelty 3.0 of 10

    GPT-4.1 predictions from fan text match self-reported experience ratings within one point 67% of the time but are biased low by one point, interpreted as measuring salient moments versus integrated overall judgment.

Pith tools