Six state-of-the-art LLMs systematically prefer Standard American English over AAE continuations, and a training-free activation steering method reduces this bias 5-20x more than prompting while preserving fluency.
Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection
14 Pith papers cite this work, alongside 52 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
representative citing papers
Evaluator-specific demonstrations with retrospective reasoning improve LLM simulation of individual preference judges by up to 9.9 points over a non-personalized base judge, while interface telemetry often degrades accuracy.
Demographic-conditioned fusion embeddings improve prediction of perspectivist social meaning interpretations by 5.9-6.5% relative macro PR-AUC over text-only baselines, with ablations confirming demographic signal.
New Ghost Annotator framework uses conformal prediction to show LLMs of different sizes and families produce labels no human annotator chose and align least with 18-30 male Sub-Saharan African annotators across content moderation datasets.
LLMs correct only 34.8% of zero-shot annotation errors via prompting, and Definition-Specific Familiarity correlates positively with performance (partial r = +0.41) while memorization metrics do not.
Large-scale statistical analysis of four harmful language datasets reveals that interactions between annotator characteristics and linguistic cues drive annotation variation, with lexical features and attitudes prominent but patterns varying by dataset.
Annotator Policy Models learn safety policies from labeling behavior alone, accurately predicting responses and revealing sources of disagreement like policy ambiguity and value pluralism.
A statistical framework decomposes human annotation outcomes into four interpretable variation sources and extends classical measurement-error models to handle both shared and individualized notions of truth.
Benchmarks for VLMs in urban perception should incorporate reliability metrics and negotiable labels, as model-human agreement co-varies with human reliability and distributional mismatches appear in a Montreal street scene study.
LLMs show split alignment with human hate speech annotations (strong on explicit attributes, inverted on evaluative ones), and attribute-based ridge regression reconstructs continuous scores with R² up to 0.71.
Automated hate speech detectors show poor alignment with heterogeneous in-group judgments on reclaimed slur usage, driven by low inter-annotator agreement and contextual features like derogatory intent.
Ethnographic study of feminist civic-tech data work argues reparative AI dataset production requires resetting accountability ties to center those harmed by current practices.
Closure of the Perspective API exposes structural dependence on a single proprietary toxicity scorer, leaving non-updatable benchmarks and irreproducible results while risking continued reliance on closed LLMs.
Simple supervision improves LLM distributional alignment with diverse population groups on three datasets, with evaluation across multiple models and prompts providing a benchmark.
citing papers explorer
-
LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering
Six state-of-the-art LLMs systematically prefer Standard American English over AAE continuations, and a training-free activation steering method reduces this bias 5-20x more than prompting while preserving fluency.
-
PERSONAJUDGE: Simulating Individual Human Preference Judgments with Evaluator-Specific Demonstration Data
Evaluator-specific demonstrations with retrospective reasoning improve LLM simulation of individual preference judges by up to 9.9 points over a non-personalized base judge, while interface telemetry often degrades accuracy.
-
Learning Perspectivist Social Meaning via Demographic-Conditioned Fusion Embeddings
Demographic-conditioned fusion embeddings improve prediction of perspectivist social meaning interpretations by 5.9-6.5% relative macro PR-AUC over text-only baselines, with ablations confirming demographic signal.
-
The Ghost Annotator: a Framework to Explore Human Label Variation in Content Moderation through Conformal Prediction
New Ghost Annotator framework uses conformal prediction to show LLMs of different sizes and families produce labels no human annotator chose and align least with 18-30 male Sub-Saharan African annotators across content moderation datasets.
-
On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance
LLMs correct only 34.8% of zero-shot annotation errors via prompting, and Definition-Specific Familiarity correlates positively with performance (partial r = +0.41) while memorization metrics do not.
-
Who and What? Using Linguistic Features and Annotator Characteristics to Analyze Annotation Variation
Large-scale statistical analysis of four harmful language datasets reveals that interactions between annotator characteristics and linguistic cues drive annotation variation, with lexical features and attitudes prominent but patterns varying by dataset.
-
Understanding Annotator Safety Policy with Interpretability
Annotator Policy Models learn safety policies from labeling behavior alone, accurately predicting responses and revealing sources of disagreement like policy ambiguity and value pluralism.
-
From Ground Truth to Measurement: A Statistical Framework for Human Labeling
A statistical framework decomposes human annotation outcomes into four interpretable variation sources and extends classical measurement-error models to handle both shared and individualized notions of truth.
-
Benchmarks for Vision-Language Models in Urban Perception Should Be Reliability-Aware and Negotiated
Benchmarks for VLMs in urban perception should incorporate reliability metrics and negotiable labels, as model-human agreement co-varies with human reliability and distributional mismatches appear in a Montreal street scene study.
-
Attribute-Based Diagnosis of LLM Alignment with Hate Speech Annotations
LLMs show split alignment with human hate speech annotations (strong on explicit attributes, inverted on evaluative ones), and attribute-based ridge regression reconstructs continuous scores with R² up to 0.71.
-
IYKYK (But AI Doesn't): Automated Content Moderation Does Not Capture Communities' Heterogeneous Attitudes Towards Reclaimed Language
Automated hate speech detectors show poor alignment with heterogeneous in-group judgments on reclaimed slur usage, driven by low inter-annotator agreement and contextual features like derogatory intent.
-
Can Data Work be Reparative?
Ethnographic study of feminist civic-tech data work argues reparative AI dataset production requires resetting accountability ties to center those harmed by current practices.
-
Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation
Closure of the Perspective API exposes structural dependence on a single proprietary toxicity scorer, leaving non-updatable benchmarks and irreproducible results while risking continued reliance on closed LLMs.
-
Improving the Distributional Alignment of LLMs using Supervision
Simple supervision improves LLM distributional alignment with diverse population groups on three datasets, with evaluation across multiple models and prompts providing a benchmark.