Large AI judges agree only weakly to moderately with human raters when ranking the harmfulness of smaller AI models' outputs, and the three small models differ in how often they produce harmful content.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
Large AI judges agree only weakly to moderately with human raters when ranking the harmfulness of smaller AI models' outputs, and the three small models differ in how often they produce harmful content.