Pith. sign in

REVIEW 2 cited by

Evaluation is all you need. Prompting Generative Large Language Models for Annotation Tasks in the Social Sciences. A Primer using Open Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.00284 v1 pith:H3SPGHBE submitted 2023-12-30 cs.CL

classification cs.CL
keywords modelsopenannotationtasksgenerativehighlightslanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper explores the use of open generative Large Language Models (LLMs) for annotation tasks in the social sciences. The study highlights the challenges associated with proprietary models, such as limited reproducibility and privacy concerns, and advocates for the adoption of open (source) models that can be operated on independent devices. Two examples of annotation tasks, sentiment analysis in tweets and identification of leisure activities in childhood aspirational essays are provided. The study evaluates the performance of different prompting strategies and models (neural-chat-7b-v3-2, Starling-LM-7B-alpha, openchat_3.5, zephyr-7b-alpha and zephyr-7b-beta). The results indicate the need for careful validation and tailored prompt engineering. The study highlights the advantages of open models for data privacy and reproducibility.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 3 citations worldwide. Full citation record

  1. Tell Me What You Know About Sexism: Expert-LLM Interaction Strategies and Co-Created Definitions for Zero-Shot Sexism Detection

    cs.CL 2025-04 conditional novelty 6.0 of 10

    Co-created expert-LLM definitions of sexism rarely beat LLM-generated definitions, but expert-written definitions perform substantially worse in zero-shot sexism detection.

  2. TextClass Benchmark: A Continuous Elo Rating of LLMs in Social Sciences

    cs.CL 2024-11 conditional novelty 5.0 of 10

    TextClass Benchmark applies a continuous Elo and Meta-Elo rating to LLMs for social-science text classification; the first snapshot covers toxicity detection in Chinese, English, German, and Russian.

Pith tools