A systematic human audit of SocialIQa, FauxPas-EAI and ToMi shows that benchmark scores are inflated or distorted by data flaws, rigid scoring, and sensitivity to phrasing.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It
A systematic human audit of SocialIQa, FauxPas-EAI and ToMi shows that benchmark scores are inflated or distorted by data flaws, rigid scoring, and sensitivity to phrasing.