Across three datasets, five language models, and six post-hoc attribution methods, explanation faithfulness, robustness, and complexity scores differ significantly between male and female inputs in a large majority of 5,040 tested configurations.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods
Across three datasets, five language models, and six post-hoc attribution methods, explanation faithfulness, robustness, and complexity scores differ significantly between male and female inputs in a large majority of 5,040 tested configurations.