REVIEW 3 major objections 6 minor 81 references
Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods
T0 review · 3 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read This paper claims that post-hoc explanation methods themselves produce explanations with significant gender disparity, independent of biases in the underlying model.
desk verdict First systematic audit of gender disparity in post-hoc explanation quality for PLMs on text; the faithfulness and complexity results are credible, but the sensitivity metric looks misimplemented, so the robustness claim needs fixing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument rests on a disparity-measurement pipeline built from male-female sentence pairs and explanation-quality metrics. Six post-hoc feature attribution methods (Gradient, Integrated Gradients, Gradient×Input, IG×Input, LIME, and SHAP) produce token-importance scores; those scores are then scored by seven metrics covering faithfulness (comprehensiveness, sufficiency, and their soft counterparts), complexity (sparsity, Gini index), and robustness (sensitivity). For each metric, the male and female score distributions are compared with the Mann-Whitney U test, and the magnitude of disparity is quantified with Cohen's $d$. The key move is running the same pipeline on models trained from scratch on a gender-balanced dataset, isolating the contribution of the explanation method from the contribution of the model.
What would settle it
Check the sensitivity implementation in the released code and verify whether the perturbation search maximizes $\|\Phi(f,y)-\Phi(f,x)\|$ or the prediction error; if it maximizes prediction error, rerun the robustness comparisons with the correct objective and see whether the gender disparities in sensitivity persist.
Extended reading notes
Core claim
The central claim is that explanation quality is not gender-neutral: feature attribution methods produce significantly different faithfulness, robustness, and complexity scores for male and female inputs, with statistically significant disparity ($p \leq 0.05$) in 3,647 of 5,040 experimental combinations and considerable effect size ($|d| \geq 0.2$) in 2,761. The effect is consistent across all six methods, with IG×Input, SHAP, and LIME showing the highest rates. The authors further claim that the disparity is not merely a consequence of biased models or data: retraining BERT and GPT-2 from scratch on the gender-balanced GECO dataset still leaves over 80% of runs with significant gender disparity. The paper concludes that post-hoc explanation methods themselves can be a source of unfairness, independent of model-level bias.
Load-bearing premise
The robustness pillar depends on the sensitivity metric being computed as defined in Eq. 6, but the paper's Appendix A.6 says the PGD attack maximizes prediction error instead of explanation change, so the reported sensitivity values may not measure explanation robustness.
Editorial extensions
If this is right
- Practitioners who audit a model's fairness by inspecting its explanations should also audit the explanations themselves, because the observed disparities can appear even when the model's predictions are not significantly biased.
- Deploying post-hoc explanation methods in high-stakes text applications without subgroup checks can produce misleading justifications for one gender, undermining trust and potentially violating transparency obligations.
- Training or fine-tuning on unbiased data is not a sufficient safeguard for explanation fairness; the attribution methods themselves need to be evaluated and, if needed, adjusted.
- Because sensitivity disparity was the most frequent on the GECO datasets, explanation robustness should be included in any fairness-oriented evaluation of attribution methods.
- Larger models such as RoBERTa reduce but do not eliminate the disparities, so model scale alone is not a remedy.
Reading between the lines
- If the attribution methods are the source of disparity, then model-level debiasing alone will not fix explanation unfairness; explanation-evaluation suites should include per-subgroup score distributions as a standard check.
- The appendix's discrepancy between the stated sensitivity objective and the PGD implementation used (prediction error instead of explanation change) suggests that the robustness results should be re-measured before being relied on, while the faithfulness and complexity findings do not depend on that step.
- A testable extension is to vary the perturbation objective and radius in the sensitivity metric to see whether gender disparity in robustness changes once the explanation change is truly maximized, and to compare hard versus soft token removal in faithfulness.
- The same disparity-measurement pipeline could be applied to non-binary gender markers or to race- and age-related cues in text, though the synthetic data generation would need to be adapted to avoid the corpus-design pitfalls the paper acknowledges.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates whether post-hoc feature attribution methods (Gradient, Gradient×Input, Integrated Gradients, IG×Input, LIME, SHAP) show gender disparities in explanation faithfulness, robustness, and complexity when explaining language model predictions. Across GECO, Stereotypes, and COMPAS text datasets and five transformer models, the authors compute seven evaluation metrics, test significance with the Mann-Whitney U test, and quantify effect sizes with Cohen's d. They report that 72.4% of 5,040 dataset-model-explainer-metric-seed combinations show statistically significant disparity, that disparities persist when models are trained from scratch on GECO, and they conclude that explanation methods themselves contribute to gender bias. The paper also discusses implications for practitioners and regulators, connecting explanation fairness to frameworks such as the EU AI Act.
Significance. The paper addresses an important and underexplored question: whether post-hoc explanation methods themselves produce systematically different explanation quality across gender subgroups. The empirical scope is substantial: six explainers, five transformer models, three datasets, seven metrics, and five seeds, plus from-scratch training experiments. The authors use established metrics and statistical tests from prior literature, report effect sizes in addition to p-values, and make a clear practical case for auditing explanation fairness. If the sensitivity implementation is corrected and the causal attribution is appropriately tempered, the study would be the first systematic NLP demonstration of explanation-level gender disparity and would provide a valuable benchmark for future XAI fairness work. The current manuscript, however, does not yet support the robustness leg of the central claim and overstates the causal role of explanation methods.
major comments (3)
- [Appendix A.6, Eq. (6)] The sensitivity metric is defined as the worst-case relative change in the explanation under an input perturbation, max over y with ||x-y||<=r of ||Phi(f,y)-Phi(f,x)|| / ||Phi(f,x)||, but the implementation is described as a PGD attack that 'perturbs the input in the direction of the gradient maximizing the prediction error.' This attack maximizes prediction loss, not the explanation-distance term in Eq. (6). Unless the two objectives are shown to coincide, which the manuscript does not demonstrate, the reported Sensitivity values measure a different quantity: the explanation change caused by a prediction-error attack. Because sensitivity is the only robustness metric, the robustness pillar of the central claim ('faithfulness, robustness, and complexity') is unsupported as reported. Please re-run the attack with the explanation-distance objective (or provide evidence that the two objectives yield equivalent worst-case perturbations) and regenerate Tables 3-5 and the affected appendix tables.
- [Section 4.4 and Section 5] The disparity analysis uses the Mann-Whitney U test at p<=0.05 for each of the 5,040 dataset-model-explainer-metric-seed combinations without any correction for multiple comparisons. With the COMPAS dataset's large male subset (n=4,997), even negligible effect sizes can become statistically significant, so the headline figure of 72.4% significant disparities may be inflated. Please report multiple-testing-corrected p-values (for example, Benjamini-Hochberg within each metric/model/dataset family or across all tests) and show how many disparities remain significant after correction; alternatively, justify why uncorrected tests are appropriate for this descriptive audit.
- [Section 5.4 and Section 4.1] The 'training from scratch on an unbiased dataset' experiments use GECO, where the classification label is the gender expressed in the sentence. Because the model is trained to predict gender, it must rely on gender-marking tokens, and the male/female inputs differ exactly in those tokens. Persistent disparity in explanation metrics under this condition is therefore not sufficient to conclude that 'explanation methods themselves can contribute to these disparities' (Section 5.4), because token-level differences between male and female inputs are a confound. The experiment rules out biased pre-training data, but it does not rule out task-induced or token-induced disparity. Please add a control condition in which gender is task-irrelevant, match male/female inputs on token-level statistics, or rephrase the conclusion to the weaker claim that disparities persist in the absence of biased training labels.
minor comments (6)
- [Section 4.2 and References] The citation for FairBERTa is [31], but reference [31] is the TinyBERT distillation paper; TinyBERT is cited as [55], which is the FairBERTa/perturbation-augmentation paper. Please swap the citations.
- [Tables 3 and 4 captions] The captions state that cell colors indicate 'which gender has better evaluation scores,' but for sensitivity, sparsity, and sufficiency lower values are preferred; Section 5 correctly says that colors indicate higher scores. Please align the captions with the text.
- [Appendix E, Table 7] Values such as '0.9910.001' are missing delimiters and plus-minus signs, making the TPR/TNR/APD entries unreadable; please reformat with proper separators.
- [Footnote 2 and Appendix B] The paper says code and datasets are released on GitHub and made public, but it provides no repository URL or version. Please add a formal availability statement with a link.
- [Appendix A.6] The normalization by ||Phi(f,x)|| in Eq. (6) can be unstable when an explanation vector is near zero; this may explain the very large and highly variable Cohen's d values in Tables 8-11 (for example, FairBERTa soft comprehensiveness d = 17.86 ± 31.01 on GECO-ALL). Please consider a regularized normalization or report unnormalized changes as a robustness check.
- [Throughout] There are several typos, including 'thecomprehensiveness' in Section 5.2.1, 'run on Stereotypes' in Section 5.2.3, 'scoresobtained' in Figure 2's caption, and 'Table in 3' in Section 5.4. A careful proofreading pass is needed.
Circularity Check
No circularity: the evaluation metrics, disparity tests, and controls are externally defined and empirically applied; the paper's self-citations are not load-bearing.
full rationale
This paper is an empirical audit rather than a derivation, so the circularity tests rarely apply. The metric definitions in Appendix A (comprehensiveness, sufficiency, soft variants, sparsity, Gini index, sensitivity) are taken from prior literature and from external implementations such as ferret; none is defined in terms of the male/female outcome that is later tested. Disparity is measured after the fact with the Mann-Whitney U test and Cohen's d, and no parameter is fitted to the disparity result: the pipeline in Algorithm 1 computes scores per input and then compares the two score lists, so there is no fitted-input-called-prediction step. The from-scratch GECO experiment is a designed control, described as initializing models randomly and training them on GECO-ALL or GECO-SUBJ; the claim that explanation methods themselves contribute to disparities is an inference drawn from observing disparity under that control, not a conclusion built into the metric definitions. The self-citations that appear (Inseq [59], the text summarization survey [20], and attention mechanisms [38]) are contextual examples or future-work pointers, and they are not used to justify the central empirical claim. The skeptical concern about Appendix A.6, where the PGD attack is described as perturbing the input in the direction of the gradient maximizing prediction error rather than maximizing the Eq. 6 explanation-change objective, is a measurement-validity issue about whether the reported sensitivity values reflect the defined quantity; it does not make the reported quantity equivalent to its input by construction. No circular step of any enumerated kind is present.
Assumptions & free parameters
free parameters (3)
- sparsity threshold tau =
0.1
- sensitivity perturbation radius r =
not reported
- PGD attack hyperparameters =
not reported
assumptions (4)
- domain assumption Mann-Whitney U is appropriate for comparing male and female explanation score distributions
- domain assumption The evaluation metrics measure the intended explanation properties (faithfulness, complexity, robustness)
- domain assumption GECO is an unbiased dataset and training on it from scratch isolates the effect of pre-training data
- domain assumption Randomly initialized BERT/GPT-2 can be trained on 3,220 sentences to yield meaningful classifiers for explanation evaluation
Cite this review
Pith. "Pith review of Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods." pith.science (2026). https://pith.science/paper/GCREUDZO
@misc{pith2026250501198,
author = {Pith},
title = {Pith review of: Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods},
year = {2026},
howpublished = {\url{https://pith.science/paper/GCREUDZO}},
note = {Machine review of arXiv:2505.01198}
}
read the original abstract
While research on applications and evaluations of explanation methods continues to expand, fairness of the explanation methods concerning disparities in their performance across subgroups remains an often overlooked aspect. In this paper, we address this gap by showing that, across three tasks and five language models, widely used post-hoc feature attribution methods exhibit significant gender disparity with respect to their faithfulness, robustness, and complexity. These disparities persist even when the models are pre-trained or fine-tuned on particularly unbiased datasets, indicating that the disparities we observe are not merely consequences of biased training data. Our results highlight the importance of addressing disparities in explanations when developing and applying explainability methods, as these can lead to biased outcomes against certain subgroups, with particularly critical implications in high-stakes contexts. Furthermore, our findings underscore the importance of incorporating the fairness of explanations, alongside overall model fairness and explainability, as a requirement in regulatory frameworks.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Definition of STAKEHOLDERS
2025. Definition of STAKEHOLDERS. https://www.merriam-webster.com/dictionary/stakeholders
2025
- [2]
-
[3]
Julia Angwin, Jeff Larson, Lauren Kirchner, and Surya Mattu. 2016. Machine bias. https://www.propublica.org/article/machine- bias-risk-assessments-in-criminal-sentencing
work page 2016
-
[4]
AI Anthropic. 2024. The claude 3 model family: Opus, sonnet, haiku. Claude-3 Model Card 1 (2024)
2024
-
[5]
Leila Arras, Ahmed Osman, and Wojciech Samek. 2022. CLEVR-XAI: A benchmark dataset for the ground truth evaluation of neural network explanations. Information Fusion 81 (2022), 14–40. doi: 10.1016/j.inffus.2021.11.008
-
[6]
Giuseppe Attanasio, Eliana Pastor, Chiara Di Bonaventura, and Debora Nozza. 2023. ferret: a Framework for Benchmarking Explainers on Transformers. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations . Association for Computational Linguistics
work page 2023
-
[7]
Aparna Balagopalan, Haoran Zhang, Kimia Hamidieh, Thomas Hartvigsen, Frank Rudzicz, and Marzyeh Ghassemi. 2022. The Road to Explainability is Paved with Bias: Measuring the Fairness of Explanations. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (Seoul, Republic of Korea) (FAccT ’22). Association for Computing Mach...
arXiv 2022
-
[8]
Esma Balkir, Svetlana Kiritchenko, Isar Nejadgholi, and Kathleen Fraser. 2022. Challenges in Applying Explainability Methods to Improve the Fairness of NLP Models. In Proceedings of the 2nd Workshop on Trustworthy Natural Language Processing (TrustNLP 2022) , Apurv Verma, Yada Pruksachatkun, Kai-Wei Chang, Aram Galstyan, Jwala Dhamala, and Yang Trista Cao...
Show all 81 references
-
[9]
Milan Bhan, Jean-Noel Vittaut, Nicolas Chesneau, and Marie-Jeanne Lesot. 2024. Self-AMPLIFY: Improving Small Language Models with Self Post Hoc Explanations. arXiv preprint arXiv:2402.12038 (2024)
2024 arXiv
-
[10]
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020. Language (Technology) is Power: A Critical Survey of “Bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, ...
2020 doi
-
[11]
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. 2021. Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 1...
2021
-
[12]
Stephanie Brandl, Emanuele Bugliarello, and Ilias Chalkidis. 2024. On the Interplay between Fairness and Explainability. In Proceedings of the 4th Workshop on Trustworthy Natural Language Processing (TrustNLP 2024) , Anaelia Ovalle, Kai-Wei Chang, Yang Trista Cao, Ninareh Mehr...
2024 doi
-
[13]
George Chrysostomou and Nikolaos Aletras. 2022. An Empirical Study on Explanations in Out-of-Domain Settings. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Smaranda Muresan, Preslav Nakov, and Aline Villavi...
2022 doi
-
[14]
Council of European Union. 2024. Council regulation (EU) no 2024/1689. https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024R1689
2024
-
[15]
Bach, and Himabindu Lakkaraju
Jessica Dai, Sohini Upadhyay, Ulrich Aivodji, Stephen H. Bach, and Himabindu Lakkaraju. 2022. Fairness via Explanation Quality: Evaluating Disparities in the Quality of Post hoc Explanations. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society (Oxford, Un...
2022
-
[16]
Marina Danilevsky, Kun Qian, Ranit Aharonov, Yannis Katsis, Ban Kawas, and Prithviraj Sen. 2020. A Survey of the State of Explainable AI for Natural Language Processing. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Lingu...
2020 doi
-
[17]
Björn Deiseroth, Mayukh Deb, Samuel Weinbach, Manuel Brack, Patrick Schramowski, and Kristian Kersting. 2023. AtMan: Understanding Transformer Predictions Through Memory Efficient Attention Manipulation. arXiv:2301.08110 [cs.LG]https://arxiv.org/abs/2301.08110
2023 arXiv
-
[18]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human...
2019
-
[19]
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2020. ERASER: A Benchmark to Evaluate Rationalized NLP Models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan J...
2020 doi
-
[20]
Mahdi Dhaini, Ege Erdogan, Smarth Bakshi, and Gjergji Kasneci. 2024. Explainability Meets Text Summarization: A Survey. In Proceedings of the 17th International Natural Language Generation Conference , Saad Mahamood, Nguyen Le Minh, and Daphne Ippolito (Eds.). Association for ...
2024
-
[21]
Xi Fang, Weijie Xu, Fiona Anting Tan, Jiani Zhang, Ziqing Hu, Yanjun Qi, Scott Nickleach, Diego Socolinsky, Srinivasan Sengamedu, and Christos Faloutsos. 2024. Large language models on tabular data–a survey. arXiv e-prints (2024), arXiv–2402
2024
-
[22]
Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K Ahmed
-
[23]
Amirata Ghorbani, Abubakar Abid, and James Zou. 2019. Interpretation of Neural Networks Is Fragile. Proceedings of the AAAI Conference on Artificial Intelligence 33, 01 (Jul. 2019), 3681–3688. doi:10.1609/aaai.v33i01.33013681
2019 doi
-
[24]
Jennifer Hsia, Danish Pruthi, Aarti Singh, and Zachary Lipton. 2024. Goodhart‘s Law Applies to NLP‘s Explanation Benchmarks. In Findings of the Association for Computational Linguistics: EACL 2024 , Yvette Graham and Matthew Purver (Eds.). Association for Computational Linguis...
2024
-
[25]
Marco Huber, Meiling Fang, Fadi Boutros, and Naser Damer. 2023. Are Explainability Tools Gender Biased? A Case Study on Face Presentation Attack Detection. In 2023 31st European Signal Processing Conference (EUSIPCO) . 945–949. doi:10.23919/EUSIPCO58844.2023.10289865
2023
-
[26]
Alon Jacovi. 2023. Trends in Explainable AI (XAI) Literature. arXiv:2301.05433 [cs.AI] https://arxiv.org/abs/2301.05433
2023 arXiv
-
[27]
Alon Jacovi and Yoav Goldberg. 2020. Towards Faithfully Interpretable NLP Systems: How Should We Define and Evaluate Faithfulness?. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel...
2020 doi
-
[28]
Sarthak Jain and Byron C. Wallace. 2019. Attention is not Explanation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , Jill Burstein, Christy D...
2019 doi
-
[29]
Sophie Jentzsch and Cigdem Turan. 2022. Gender Bias in BERT-Measuring and Analysing Biases through Sentiment Rating in a Realistic Downstream Classification Task. In Proceedings of the 4th Workshop on Gender Bias in Natural Language Processing (GeBNLP) . 184–199
2022
-
[30]
Sérgio Jesus, Catarina Belém, Vladimir Balayan, João Bento, Pedro Saleiro, Pedro Bizarro, and João Gama. 2021. How can I choose an explainer? An Application-grounded Evaluation of Post-hoc Explanations. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and...
2021
-
[31]
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020. TinyBERT: Distilling BERT for Natural Language Understanding. In Findings of the Association for Computational Linguistics: EMNLP 2020 , Trevor Cohn, Yulan He, and Yang Liu (E...
2020 doi
-
[32]
Shreya Johri, Jaehwan Jeong, Benjamin A Tran, Daniel I Schlessinger, Shannon Wongvibulsin, Leandra A Barnes, Hong-Yu Zhou, Zhuo Ran Cai, Eliezer M Van Allen, David Kim, et al. 2025. An evaluation framework for clinical use of large language models in patient interaction tasks....
2025 doi
-
[33]
Neema Kotonya and Francesca Toni. 2020. Explainable Automated Fact-Checking: A Survey. In Proceedings of the 28th International Conference on Computational Linguistics, Donia Scott, Nuria Bel, and Chengqing Zong (Eds.). International Committee on Computational Linguistics, Bar...
2020 doi
-
[34]
Satyapriya Krishna, Tessa Han, Alex Gu, Steven Wu, Shahin Jabbari, and Himabindu Lakkaraju. 2024. The Disagreement Problem in Explainable Machine Learning: A Practitioner’s Perspective. Transactions on Machine Learning Research (2024). https://openreview.net/forum?id= jESY2WTZCe
2024
-
[35]
Satyapriya Krishna, Jiaqi Ma, Dylan Slack, Asma Ghandeharioun, Sameer Singh, and Himabindu Lakkaraju. 2024. Post hoc explanations of language models can improve language models. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New O...
2024
-
[36]
How do I fool you?
Himabindu Lakkaraju and Osbert Bastani. 2020. "How do I fool you?": Manipulating User Trust via Misleading Black Box Explanations. InProceedings of the AAAI/ACM Conference on AI, Ethics, and Society (New York, NY, USA) (AIES ’20). Association for Computing Machinery, New York,...
2020
-
[37]
Markus Langer, Daniel Oster, Timo Speith, Holger Hermanns, Lena Kästner, Eva Schmidt, Andreas Sesing, and Kevin Baum. 2021. What do we want from Explainable Artificial Intelligence (XAI)? – A stakeholder perspective on XAI and a conceptual model guiding interdisciplinary XAI r...
2021
-
[38]
Tobias Leemann, Alina Fastowski, Felix Pfeiffer, and Gjergji Kasneci. 2025. Attention Mechanisms Don’t Learn Additive Models: Rethinking Feature Importance for Transformers. arXiv:2405.13536 [cs.LG] https://arxiv.org/abs/2405.13536 Gender Bias in Explainability: Investigating ...
2025 arXiv
-
[39]
Xuhong Li, Mengnan Du, Jiamin Chen, Yekun Chai, Himabindu Lakkaraju, and Haoyi Xiong. 2023. M4: A Unified XAI Benchmark for Faithfulness Evaluation of Feature Attribution Methods across Metrics, Modalities and Models. InAdvances in Neural Information Processing Systems, A. Oh,...
2023
-
[40]
Hui Liu, Qingyu Yin, and William Yang Wang. 2019. Towards Explainable NLP: A Generative Explanation Framework for Text Classification. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , Anna Korhonen, David Traum, and Lluís Màrquez (Ed...
2019 doi
-
[41]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR abs/1907.11692 (2019). arXiv:1907.11692 http://arxiv.org/abs/1907. 11692
2019 arXiv
-
[42]
Luca Longo, Mario Brcic, Federico Cabitza, Jaesik Choi, Roberto Confalonieri, Javier Del Ser, Riccardo Guidotti, Yoichi Hayashi, Francisco Herrera, Andreas Holzinger, Richard Jiang, Hassan Khosravi, Freddy Lecue, Gianclaudio Malgieri, Andrés Páez, Wojciech Samek, Johannes Schn...
2024
-
[43]
I Loshchilov. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101 (2017)
2017 arXiv
-
[44]
Scott M Lundberg and Su-In Lee. 2017. A unified approach to interpreting model predictions. Advances in neural information processing systems 30 (2017)
2017
-
[45]
Qing Lyu, Marianna Apidianaki, and Chris Callison-Burch. 2024. Towards Faithful Model Explanation in NLP: A Survey. Computational Linguistics 50, 2 (June 2024), 657–723. doi:10.1162/coli_a_00511
2024 doi
-
[46]
Aleksander Madry. 2017. Towards deep learning models resistant to adversarial attacks. arXiv preprint arXiv:1706.06083 (2017)
2017 arXiv
-
[47]
Andreas Madsen, Himabindu Lakkaraju, Siva Reddy, and Sarath Chandar. 2024. Interpretability Needs a New Paradigm. arXiv:2405.05386 [cs.LG] https://arxiv.org/abs/2405.05386
2024 arXiv
-
[48]
Andreas Madsen, Siva Reddy, and Sarath Chandar. 2022. Post-hoc Interpretability for Neural NLP: A Survey. ACM Comput. Surv. 55, 8, Article 155 (Dec. 2022), 42 pages. doi: 10.1145/3546577
2022 doi
-
[49]
Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. 2021. HateXplain: A Benchmark Dataset for Explainable Hate Speech Detection. Proceedings of the AAAI Conference on Artificial Intelligence 35, 17 (May 2021), 14867–14875. doi:10.1...
2021 doi
-
[50]
Vishwali Mhasawade, Salman Rahman, Zoé Haskell-Craig, and Rumi Chunara. 2024. Understanding Disparities in Post Hoc Machine Learning Explanation. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (Rio de Janeiro, Brazil) (FAccT ’24). Assoc...
2024
-
[51]
Katelyn Morrison, Philipp Spitzer, Violet Turri, Michelle Feng, Niklas Kühl, and Adam Perer. 2024. The Impact of Imperfect XAI on Human-AI Decision-Making. Proc. ACM Hum.-Comput. Interact. 8, CSCW1, Article 183 (April 2024), 39 pages. doi:10.1145/3641022
2024 doi
-
[52]
Edoardo Mosca, Ferenc Szigeti, Stella Tragianni, Daniel Gallagher, and Georg Groh. 2022. SHAP-Based Explanation Methods: A Review for NLP Interpretability. In Proceedings of the 29th International Conference on Computational Linguistics , Nicoletta Calzolari, Chu-Ren Huang, Ha...
2022
-
[53]
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020. CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , Bonnie Webber, T...
2020 doi
-
[54]
Meike Nauta, Jan Trienes, Shreyasi Pathak, Elisa Nguyen, Michelle Peters, Yasmin Schmitt, Jörg Schlötterer, Maurice van Keulen, and Christin Seifert. 2023. From Anecdotal Evidence to Quantitative Evaluation Methods: A Systematic Review on Evaluating Explainable AI. ACM Comput....
2023 doi
-
[55]
Rebecca Qian, Candace Ross, Jude Fernandes, Eric Michael Smith, Douwe Kiela, and Adina Williams. 2022. Perturbation Augmentation for Fairer NLP. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . 9496–9521
2022
-
[56]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9
2019
-
[57]
Why should i trust you?
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016. " Why should i trust you?" Explaining the predictions of any classifier. InProceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining . 1135–1144
2016
-
[58]
Marko Robnik-Šikonja and Marko Bohanec. 2018. Perturbation-Based Explanations of Prediction Models . Springer International Publishing, Cham, 159–175. doi: 10.1007/978-3-319-90403-0_9
2018 doi
-
[59]
Gabriele Sarti, Nils Feldhus, Ludwig Sickert, and Oskar van der Wal. 2023. Inseq: An Interpretability Toolkit for Sequence Generation Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), Danushka ...
2023 doi
-
[60]
Shlomo S Sawilowsky. 2009. New effect size rules of thumb. Journal of modern applied statistical methods 8 (2009), 597–599. 18 Dhaini et al
2009
-
[61]
Jakob Schoeffer, Maria De-Arteaga, and Niklas Kühl. 2024. Explanations, Fairness, and Appropriate Reliance in Human-AI Decision-Making. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Mach...
2024
-
[62]
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2013. Deep inside convolutional networks: Visualising image classification models and saliency maps. arXiv preprint arXiv:1312.6034 (2013)
2013 arXiv
-
[63]
Samuel Sithakoul, Sara Meftah, and Clément Feutry. 2024. BEExAI: Benchmark to Evaluate Explainable AI. In Explainable Artificial Intelligence, Luca Longo, Sebastian Lapuschkin, and Christin Seifert (Eds.). Springer Nature Switzerland, Cham, 445–468
2024
-
[64]
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017. Axiomatic attribution for deep networks. In International conference on machine learning . PMLR, 3319–3328
2017
-
[65]
Santosh T.y.s.s., Nina Baumgartner, Matthias Stürmer, Matthias Grabmair, and Joel Niklaus. 2024. Towards Explainability and Fairness in Swiss Judgement Prediction: Benchmarking on a Multilingual Dataset. InProceedings of the 2024 Joint International Conference on Computational...
2024
-
[66]
Josef Valvoda and Ryan Cotterell. 2024. Towards Explainability in Legal Outcome Prediction Models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , Kevin ...
2024 doi
-
[67]
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. 2020. Investigating Gender Bias in Language Models Using Causal Mediation Analysis. In Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R....
2020
-
[68]
Eric Wallace, Matt Gardner, and Sameer Singh. 2020. Interpreting Predictions of NLP Models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts , Aline Villavicencio and Benjamin Van Durme (Eds.). Association for Comput...
2020 doi
-
[69]
Rick Wilming, Artur Dox, Hjalmar Schulz, Marta Oliveira, Benedict Clark, and Stefan Haufe. 2024. GECOBench: A Gender-Controlled Text Dataset and Benchmark for Quantifying Biases in Explanations. arXiv preprint arXiv:2406.11547 (2024)
2024
-
[70]
Alice Xiang and Inioluwa Deborah Raji. 2019. On the Legal Compatibility of Fairness Definitions. arXiv:1912.00761 [cs.CY] https://arxiv. org/abs/1912.00761
2019 arXiv
-
[71]
Wenzhuo Yang, Hung Le, Tanmay Laud, Silvio Savarese, and Steven C. H. Hoi. 2022. OmniXAI: A Library for Explainable AI. arXiv:2206.01612 [cs.LG] https://arxiv.org/abs/2206.01612
2022 arXiv
-
[72]
Chih-Kuan Yeh, Cheng-Yu Hsieh, Arun Suggala, David I Inouye, and Pradeep K Ravikumar. 2019. On the (In)fidelity and Sensitivity of Explanations. In Advances in Neural Information Processing Systems , H. Wallach, H. Larochelle, A. Beygelzimer, F. d 'Alché-Buc, E. Fox, and R. Ga...
2019
-
[73]
Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du. 2024. Explainability for Large Language Models: A Survey. ACM Trans. Intell. Syst. Technol. 15, 2, Article 20 (Feb. 2024), 38 pages. doi: 10.1145/3639372
2024 doi
-
[74]
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018. Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human...
2018 doi
-
[75]
Zhixue Zhao and Nikolaos Aletras. 2023. Incorporating Attribution Importance for Improving Faithfulness Metrics. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Anna Rogers, Jordan Boyd-Graber, and Naoaki Oka...
2023 doi
-
[76]
Zhixue Zhao, George Chrysostomou, Kalina Bontcheva, and Nikolaos Aletras. 2022. On the Impact of Temporal Concept Drift on Model Explanations. In Findings of the Association for Computational Linguistics: EMNLP 2022 , Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Ass...
2022 doi
-
[77]
Gandomi, Fang Chen, and Andreas Holzinger
Jianlong Zhou, Amir H. Gandomi, Fang Chen, and Andreas Holzinger. 2021. Evaluating the Quality of Machine Learning Explanations: A Survey on Methods and Metrics. Electronics 10, 5 (2021). doi: 10.3390/electronics10050593
2021 doi
-
[78]
Julia El Zini and Mariette Awad. 2022. On the Explainability of Natural Language Processing Deep Models. ACM Comput. Surv. 55, 5, Article 103 (Dec. 2022), 31 pages. doi: 10.1145/3529755 Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods 19 A...
2022 doi
-
[80]
I was surprised to see a woman doctor articulate herself so well
and COMPAS datasets [3], as well as the synthetic Stereotypes dataset we create and make public. We use the publicly available models from Huggingface (see Table 6) running on a single NVIDIA V100 GPU. Including fine-tuning, and generating and evaluating explanations, one mode...
2024
-
[81]
significant
with initial learning rate 0.001 and a linear learning rate schedule with 500 warm-up steps. Since our tasks are binary classification tasks, we use the binary cross-entropy loss. We observed that after one epoch of fine-tuning, the models perform hardly better than random gue...
-
[2024]
Computational Linguistics (2024), 1–79
Bias and fairness in large language models: A survey. Computational Linguistics (2024), 1–79
2024
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.