LLMs become less accurate and more overconfident as human annotator agreement drops, and training on disagreement samples improves in-domain accuracy and confidence alignment.
Title resolution pending
1 Pith paper cite this work, alongside 47 external citations. Polarity classification is still indexing.
1
Pith paper citing it
47
external citations · OpenAlex
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
LLMs become less accurate and more overconfident as human annotator agreement drops, and training on disagreement samples improves in-domain accuracy and confidence alignment.