Ethical refusals in LLM responses sharply reduce user win rates in Chatbot Arena compared to technical refusals and normal answers, though detailed refusals and clearly harmful prompts reduce the penalty.
That said, interpretations of specific content categories should be made with caution
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
LLM Content Moderation and User Satisfaction: Evidence from Response Refusals in Chatbot Arena
Ethical refusals in LLM responses sharply reduce user win rates in Chatbot Arena compared to technical refusals and normal answers, though detailed refusals and clearly harmful prompts reduce the penalty.