New Ghost Annotator framework uses conformal prediction to show LLMs of different sizes and families produce labels no human annotator chose and align least with 18-30 male Sub-Saharan African annotators across content moderation datasets.
D 3 CODE : Disentangling Disagreements in Data across Cultures on Offensiveness Detection and Evaluation
2 Pith papers cite this work, alongside 4 external citations. Polarity classification is still indexing.
fields
cs.CL 2years
2026 2representative citing papers
Large-scale statistical analysis of four harmful language datasets reveals that interactions between annotator characteristics and linguistic cues drive annotation variation, with lexical features and attitudes prominent but patterns varying by dataset.
citing papers explorer
-
The Ghost Annotator: a Framework to Explore Human Label Variation in Content Moderation through Conformal Prediction
New Ghost Annotator framework uses conformal prediction to show LLMs of different sizes and families produce labels no human annotator chose and align least with 18-30 male Sub-Saharan African annotators across content moderation datasets.
-
Who and What? Using Linguistic Features and Annotator Characteristics to Analyze Annotation Variation
Large-scale statistical analysis of four harmful language datasets reveals that interactions between annotator characteristics and linguistic cues drive annotation variation, with lexical features and attitudes prominent but patterns varying by dataset.