Pith. sign in

REVIEW 4 cited by

Two Contrasting Data Annotation Paradigms for Subjective NLP Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.07475 v2 pith:CQWTLMTX submitted 2021-12-14 cs.CL

classification cs.CL
keywords annotationdataparadigmsdatasetannotatorbeliefscontrastingcreators
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Labelled data is the foundation of most natural language processing tasks. However, labelling data is difficult and there often are diverse valid beliefs about what the correct data labels should be. So far, dataset creators have acknowledged annotator subjectivity, but rarely actively managed it in the annotation process. This has led to partly-subjective datasets that fail to serve a clear downstream use. To address this issue, we propose two contrasting paradigms for data annotation. The descriptive paradigm encourages annotator subjectivity, whereas the prescriptive paradigm discourages it. Descriptive annotation allows for the surveying and modelling of different beliefs, whereas prescriptive annotation enables the training of models that consistently apply one belief. We discuss benefits and challenges in implementing both paradigms, and argue that dataset creators should explicitly aim for one or the other to facilitate the intended use of their dataset. Lastly, we conduct an annotation experiment using hate speech data that illustrates the contrast between the two paradigms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.

  2. The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making

    cs.AI 2025-06 reject novelty 6.0 of 10

    MedPerturb finds that LLMs are more sensitive to gender and style changes in clinical text, while medical students are more sensitive to LLM-generated summaries and dialogues, in triage decisions.

  3. The Generative AI Ethics Playbook

    cs.CY 2024-12 conditional novelty 4.0 of 10

    A structured playbook that collects existing guidance, checklists, and case studies to help generative AI practitioners identify and mitigate ethical harms across six lifecycle stages.

  4. An Annotated Corpus of Arabic Tweets for Hate Speech Analysis

    cs.CL 2025-05 reject novelty 3.0 of 10

    The paper presents a new 10,000-tweet Arabic hate speech corpus with seven target categories and reports transformer benchmark results, but its novelty and reliability claims are undermined by internal inconsistencies.

Pith tools