LLM-as-a-Judge systems rate ethical refusal responses more favorably than human users, a gap the paper calls moderation bias, while technical refusals do not show the same divergence.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.HC 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
AI vs. Human Judgment of Content Moderation: LLM-as-a-Judge and Ethics-Based Response Refusals
LLM-as-a-Judge systems rate ethical refusal responses more favorably than human users, a gap the paper calls moderation bias, while technical refusals do not show the same divergence.