Pith. sign in

REVIEW 3 cited by

Best Practices for Text Annotation with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.05129 v1 pith:5HLWZZIP submitted 2024-02-05 cs.CL

Best Practices for Text Annotation with Large Language Models

classification cs.CL
keywords llmsannotationpracticestextbestcriticalethicallanguage
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Large Language Models (LLMs) have ushered in a new era of text annotation, as their ease-of-use, high accuracy, and relatively low costs have meant that their use has exploded in recent months. However, the rapid growth of the field has meant that LLM-based annotation has become something of an academic Wild West: the lack of established practices and standards has led to concerns about the quality and validity of research. Researchers have warned that the ostensible simplicity of LLMs can be misleading, as they are prone to bias, misunderstandings, and unreliable results. Recognizing the transformative potential of LLMs, this paper proposes a comprehensive set of standards and best practices for their reliable, reproducible, and ethical use. These guidelines span critical areas such as model selection, prompt engineering, structured prompting, prompt stability analysis, rigorous model validation, and the consideration of ethical and legal implications. The paper emphasizes the need for a structured, directed, and formalized approach to using LLMs, aiming to ensure the integrity and robustness of text annotation practices, and advocates for a nuanced and critical engagement with LLMs in social scientific research.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. The Role of Online Forums in Developer Understanding of Privacy Law -- A Reddit Case Study

    cs.CR 2026-06 unverdicted novelty 5.0

    Survey of 223 Reddit users and qualitative analysis of 2,248 posts shows certified developers frequently use online forums for privacy law advice, identifying key challenges and credibility assessment methods.

  2. Mapping Election Toxicity on Social Media across Issue, Ideology, and Psychosocial Dimensions

    cs.SI 2026-04 unverdicted novelty 5.0

    Large-scale analysis of election tweets finds highest toxicity intensity in identity issues, harassment as the dominant harm type, partisan posts more toxic than neutral with issue-varying asymmetries, and toxic conte...

  3. Do BERT Embeddings Encode Narrative Dimensions? A Token-Level Probing Analysis of Time, Space, Causality, and Character in Fiction

    cs.CL 2026-04 unverdicted novelty 5.0

    BERT embeddings encode narrative dimensions of time, space, causality, and character at the token level, as a linear probe achieves 94% accuracy versus 47% on variance-matched random embeddings, though unsupervised cl...