REVIEW 2 cited by
BiasX: "Thinking Slow" in Toxic Content Moderation with Explanations of Implied Social Biases
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Toxicity annotators and content moderators often default to mental shortcuts when making decisions. This can lead to subtle toxicity being missed, and seemingly toxic but harmless content being over-detected. We introduce BiasX, a framework that enhances content moderation setups with free-text explanations of statements' implied social biases, and explore its effectiveness through a large-scale crowdsourced user study. We show that indeed, participants substantially benefit from explanations for correctly identifying subtly (non-)toxic content. The quality of explanations is critical: imperfect machine-generated explanations (+2.4% on hard toxic examples) help less compared to expert-written human explanations (+7.2%). Our results showcase the promise of using free-text explanations to encourage more thoughtful toxicity moderation.
Forward citations
Cited by 2 Pith papers
-
Real-World Gaps in AI Governance Research
Corporate AI safety research is dominated by pre-deployment alignment and evaluation work, while high-risk deployment topics such as medical error, misinformation, bias, behavioral design, and copyright are measured t...
-
Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails
Aegis2.0 provides a commercially usable, human-annotated safety dataset with 24 risk categories, and models trained on it with parameter-efficient methods match WildGuard and beat Llama Guard 3.
Discussion (0). Continue with ORCID to comment.