A 14-probe benchmark on 12 LLMs finds consistent gender stereotype reasoning and unbalanced character representation across models.
Pappas, Florian Tram \` e r, Hamed Hassani, and Eric Wong
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
GenderBench: Evaluation Suite for Gender Biases in LLMs
A 14-probe benchmark on 12 LLMs finds consistent gender stereotype reasoning and unbalanced character representation across models.