Pith. sign in

REVIEW 1 cited by

Prompt Design Matters for Computational Social Science Tasks but in Unpredictable Ways

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.11980 v1 pith:CJ6RQ56K submitted 2024-06-17 cs.AI cs.CY

classification cs.AIcs.CY
keywords promptaccuracycompliancedesigntasksllmsannotationschanges
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Manually annotating data for computational social science tasks can be costly, time-consuming, and emotionally draining. While recent work suggests that LLMs can perform such annotation tasks in zero-shot settings, little is known about how prompt design impacts LLMs' compliance and accuracy. We conduct a large-scale multi-prompt experiment to test how model selection (ChatGPT, PaLM2, and Falcon7b) and prompt design features (definition inclusion, output type, explanation, and prompt length) impact the compliance and accuracy of LLM-generated annotations on four CSS tasks (toxicity, sentiment, rumor stance, and news frames). Our results show that LLM compliance and accuracy are highly prompt-dependent. For instance, prompting for numerical scores instead of labels reduces all LLMs' compliance and accuracy. The overall best prompting setup is task-dependent, and minor prompt changes can cause large changes in the distribution of generated labels. By showing that prompt design significantly impacts the quality and distribution of LLM-generated annotations, this work serves as both a warning and practical guide for researchers and practitioners.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Integrating LLMs and Digital Twins for Adaptive Multi-Robot Task Allocation in Construction

    cs.RO 2025-06 conditional novelty 5.0 of 10

    The paper integrates digital twins, integer programming, and LLMs so that natural-language site updates can automatically adapt multi-robot construction task allocation.

Pith tools