Pith. sign in

REVIEW 2 cited by

BiasX: "Thinking Slow" in Toxic Content Moderation with Explanations of Implied Social Biases

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.13589 v1 pith:4OW4S2K7 submitted 2023-05-23 cs.CL

classification cs.CL
keywords explanationscontenttoxicmoderationtoxicitybiasesbiasxfree-text
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Toxicity annotators and content moderators often default to mental shortcuts when making decisions. This can lead to subtle toxicity being missed, and seemingly toxic but harmless content being over-detected. We introduce BiasX, a framework that enhances content moderation setups with free-text explanations of statements' implied social biases, and explore its effectiveness through a large-scale crowdsourced user study. We show that indeed, participants substantially benefit from explanations for correctly identifying subtly (non-)toxic content. The quality of explanations is critical: imperfect machine-generated explanations (+2.4% on hard toxic examples) help less compared to expert-written human explanations (+7.2%). Our results showcase the promise of using free-text explanations to encourage more thoughtful toxicity moderation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Real-World Gaps in AI Governance Research

    cs.AI 2025-04 conditional novelty 6.0 of 10

    Corporate AI safety research is dominated by pre-deployment alignment and evaluation work, while high-risk deployment topics such as medical error, misinformation, bias, behavioral design, and copyright are measured t...

  2. Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Aegis2.0 provides a commercially usable, human-annotated safety dataset with 24 risk categories, and models trained on it with parameter-efficient methods match WildGuard and beat Llama Guard 3.

Pith tools