Evaluator-specific demonstrations with retrospective reasoning improve LLM simulation of individual preference judges by up to 9.9 points over a non-personalized base judge, while interface telemetry often degrades accuracy.
We need to consider disagreement in evaluation
6 Pith papers cite this work, alongside 71 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 6roles
background 1polarities
background 1representative citing papers
Demographic-conditioned fusion embeddings improve prediction of perspectivist social meaning interpretations by 5.9-6.5% relative macro PR-AUC over text-only baselines, with ablations confirming demographic signal.
Large-scale statistical analysis of four harmful language datasets reveals that interactions between annotator characteristics and linguistic cues drive annotation variation, with lexical features and attitudes prominent but patterns varying by dataset.
Disagreement in health-literacy annotations is driven by conceptual task difficulty rather than annotator differences, with social effects varying or reversing by agreement level, making perspectivist modeling necessary.
Automated hate speech detectors show poor alignment with heterogeneous in-group judgments on reclaimed slur usage, driven by low inter-annotator agreement and contextual features like derogatory intent.
citing papers explorer
-
PERSONAJUDGE: Simulating Individual Human Preference Judgments with Evaluator-Specific Demonstration Data
Evaluator-specific demonstrations with retrospective reasoning improve LLM simulation of individual preference judges by up to 9.9 points over a non-personalized base judge, while interface telemetry often degrades accuracy.
-
Learning Perspectivist Social Meaning via Demographic-Conditioned Fusion Embeddings
Demographic-conditioned fusion embeddings improve prediction of perspectivist social meaning interpretations by 5.9-6.5% relative macro PR-AUC over text-only baselines, with ablations confirming demographic signal.
-
Who and What? Using Linguistic Features and Annotator Characteristics to Analyze Annotation Variation
Large-scale statistical analysis of four harmful language datasets reveals that interactions between annotator characteristics and linguistic cues drive annotation variation, with lexical features and attitudes prominent but patterns varying by dataset.
-
Structured Disagreement in Health-Literacy Annotation: Epistemic Stability, Conceptual Difficulty, and Agreement-Stratified Inference
Disagreement in health-literacy annotations is driven by conceptual task difficulty rather than annotator differences, with social effects varying or reversing by agreement level, making perspectivist modeling necessary.
-
IYKYK (But AI Doesn't): Automated Content Moderation Does Not Capture Communities' Heterogeneous Attitudes Towards Reclaimed Language
Automated hate speech detectors show poor alignment with heterogeneous in-group judgments on reclaimed slur usage, driven by low inter-annotator agreement and contextual features like derogatory intent.
- Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling