REVIEW 4 major objections 5 minor 49 references
Explanations as Bias Detectors: A Critical Study of Local Post-hoc XAI Methods for Fairness Exploration
T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Aggregated local explanations can serve as bias detectors, flagging unequal treatment across demographic groups when standard fairness metrics are violated.
desk verdict A useful empirical study of XAI-as-bias-detector whose central claim lacks a negative control: every model tested is unfair, so the detector signal is never checked when fairness holds. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism is the aggregated group-level explanation profile. For LIME and SHAP, each instance in a group and outcome subset contributes a signed feature-attribution vector; the pipeline averages these vectors per group, and the difference between protected and non-protected group averages is the fairness signal. For DiCE, the pipeline aggregates counterfactuals by the percentage of feature changes and by the mean Euclidean distance between factual and counterfactual instances (the Burden metric). The AOPC curve is used as a check on whether the underlying per-instance feature rankings are meaningful.
What would settle it
Construct a classifier on a dataset in which the protected attribute is correlated with the outcome but the classifier is trained under a fairness constraint so that demographic parity, equal opportunity, and equalized odds are satisfied; if the pipeline still shows large unequal protected-attribute contributions across groups, the claimed correspondence would be refuted. Conversely, on a synthetic biased model with no outcome disparity, the explanation signal should be absent if it is a reliable detector.
Extended reading notes
Core claim
The central claim is a correspondence between distributive fairness and explanation-based procedural signals: whenever demographic parity, equal opportunity, or equalized odds are violated, group-level feature attribution also shows unequal contributions, typically with the protected attribute positively favoring the non-protected group and negatively affecting the protected group. In the income datasets used (Adult, AdultCA, AdultLA), LIME and SHAP average contributions for sex are negative for women and positive for men across positive, true-positive, and false-positive subsets, with the largest gaps in Louisiana; DiCE shows women and Black individuals needing to change more features, often including sex or race themselves. When the protected attribute is removed from the model, accuracy drops only slightly and fairness violations persist, with contributions shifting to correlated proxies. The paper reads this as evidence that aggregated local explanations can reveal both direct and indirect discrimination, provided the pipeline steps are chosen deliberately.
Load-bearing premise
The pipeline explains only 100 instances per demographic group and outcome category (50 for AdultLA) and assumes these are representative of the full groups, but no sampling procedure, variance estimate, or confidence intervals are reported, so the group-level attribution differences could be artifacts of the small samples.
Editorial extensions
If this is right
- Auditors can run this pipeline on an already-deployed black-box model to get a procedural-fairness read without retraining or accessing the training data.
- Removing a protected attribute from the feature set will not by itself erase the fairness signal; explanations can trace where the bias moves, pointing auditors to proxy features.
- Choosing the right aggregation is essential: absolute-value aggregations can hide oppositional bias, while signed aggregation or counterfactual burden shows which group is pushed toward unfavorable outcomes.
- Agreement among LIME, SHAP, and DiCE on the same group-level pattern strengthens the case that the observed signal is genuine rather than an artifact of one explainer.
- The AOPC result suggests the feature rankings used for these comparisons are informative, supporting the use of the pipeline in practice.
Reading between the lines
- The correspondence could be made quantitative: one could train a meta-classifier that predicts distributive fairness violations from explanation profiles, yielding an automatic fairness audit tool for settings where outcome labels are unavailable.
- The reliance on small explanation samples points to adding variance estimates and statistical tests for group attribution differences, which would turn the qualitative pattern into a rigorous audit procedure.
- The findings suggest a caution for mitigation: removing protected attributes from a model is insufficient, and explanation-based auditing should be paired with interventions on correlated features.
- The approach could extend to multiple protected attributes and intersectional groups, though pairwise comparisons would then need correction for multiple testing.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a general pipeline for using local post-hoc explanation methods (LIME, SHAP, DiCE) as bias detectors: it generates individual explanations for instances in demographic groups, aggregates them per group, and compares aggregated contributions across groups to derive fairness-related signals. It addresses four research questions: the relationship between distributive fairness metrics and explanation-based 'procedural' signals (RQ1), the effect of removing the protected attribute (RQ2), the impact of alternative aggregation strategies (RQ3), and the trustworthiness of the resulting explanations via AOPC curves (RQ4). Experiments are run on Adult, AdultCA, AdultLA, and COMPAS with Random Forest and XGBoost models. The paper concludes that when distributive fairness is violated, aggregated explanations show similar signs of procedural unfairness, such as unequal contributions of protected attributes or disproportionate use of irrelevant features, and that explanation-based signals are broadly consistent across methods, while cautioning that pipeline design choices matter.
Significance. If the claimed relationship between distributive unfairness and explanation-based procedural signals were established, the pipeline would offer a practical auditing tool for deployed models where outcome labels or fairness metrics are unavailable. The study is empirically broad and thoughtfully designed in several respects: it aligns explanation targets with the definitions of Demographic Parity, Equal Opportunity, and Equalized Odds; it selects one method per explanation family; it includes two model families and multiple datasets; and it explicitly confronts the aggregation-choice problem, showing with Figure 8 that sign-preserving versus absolute-value aggregation yields different fairness conclusions. The critical risk is that the central implication (Section 6) is tested only on models that all violate distributive fairness, so the evidence cannot distinguish the claimed detector from one that always raises an alarm. A second, compounding risk is that group-level aggregated contributions are computed on small samples of unexplained size and without variance estimates.
major comments (4)
- [Section 6 / Tables 1, 8, 9] The central claim, stated in the conclusion as 'when distributive fairness is violated we get similar signs of procedural unfairness, such as unequal contributions of protected attributes across groups or disproportionate use of irrelevant features,' is an implication whose antecedent is true in every experimental condition. In Tables 1, 8, and 9, all reported differences in PR, TPR, and FPR are statistically significant for every dataset, model, and protected attribute, so there is no condition in which Demographic Parity, Equal Opportunity, or Equalized Odds is satisfied. Consequently, the observed association is established only on positive examples, and a detector that always returns 'unfair' would reproduce the pattern. To support the implication, the paper needs at least one negative control: for example, a model trained with a preprocessing or postprocessing debiasing intervention that achieves (approximately) satisfied fairness metrics, or a synthetic dataset constructed so that the fairness metrics hold, with the prediction that the procedural signals disappear or weaken under such a condition. If such an experiment is not feasible, the conclusion should be explicitly weakened to an observation about the tested unfair models rather than an implication linking distributive and procedural fairness.
- [Section 4.1 / Figures 2-5, Tables 2 and 6] The paper generates explanations for 'a representative number of instances,' namely 100 instances per demographic group and outcome category (50 for AdultLA), but does not specify the sampling procedure, the population from which the sample is drawn, or any variance estimate. All group-level comparisons of feature contributions are point estimates without confidence intervals or standard errors. If the sampled instances are not representative of the full groups, the observed group differences in feature contributions, and hence the fairness signals derived from them, could be sampling artifacts. The authors should state how the instances were sampled (e.g., uniform random, stratified by outcome), justify the sample size, and report bootstrap confidence intervals or at least standard errors for the aggregated contributions in Figures 2-5 and Tables 2 and 6.
- [Section 4.4 / Figure 9] The AOPC trustworthiness evaluation is performed only on the Adult dataset with the Random Forest model (Figure 9), yet the conclusion in Section 6 states, 'using the AOPC curve, we conclude that the different explanation methods show consistency' and suggests the explanations can be trusted 'to some extent.' A single dataset-model combination is too narrow a basis for a trustworthiness conclusion in a paper that spans four datasets and two model families. The AOPC analysis should be extended to at least one additional dataset or model family, or the trustworthiness conclusion should be explicitly restricted to the Adult/Random Forest setting.
- [Section 4.4, paragraph 1] The AOPC evaluation is mildly self-referential: the perturbation order is derived from the same explanation ranking that is being validated. The paper cites criticism of AOPC (references [24,25]) but does not discuss how the circularity between the perturbation ordering and the evaluated explanation ranking affects the interpretation of Figure 9. A short discussion of this limitation, and of why the comparison against a random ranking partially addresses it, would make the trustworthiness argument more careful.
minor comments (5)
- [Table 4] The AdultLA FN Black entry (22.98) and AdultLA TN Female entry (20.64) are an order of magnitude larger than all other values in the same column and are likely typographical errors or anomalies that need explanation; if they are genuine, the paper should comment on why these subgroups exhibit such extreme distances.
- [Figures 13 and 14 captions] Figure 13 is captioned 'when the protected attribute race is removed' but the surrounding text in Appendix A.3 describes Figure 13 as showing removal of the sex attribute; the captions for Figures 13 and 14 appear to be swapped.
- [Section 3, aggregation formulas] The two aggregation formulas near the description of RQ3 are both labeled I_abs, although the second formula sums signed contributions without absolute values; using distinct names (e.g., I_abs and I_mean) or explicit descriptions would prevent confusion.
- [Tables 2 and 6 captions] For the DiCE columns, the reported values are percentages of feature changes, but the captions do not state the unit; adding 'percent of feature changes' for DiCE columns would increase clarity.
- [Section 3, Burden metric] The Burden metric definition uses a distance function c(x_i, x'_i) described only as 'some distance metric such as the Euclidean distance'; since the metric is later used in Table 4, the specific distance used for the experiments should be stated explicitly in Section 4.1 or in the definition.
Circularity Check
No circularity: the central claim is an empirical association, not a derivation from its inputs.
full rationale
This paper makes no formal derivation; its central claim is an observed empirical association between distributive fairness violations and aggregated explanation signals. The fairness metrics (Eqs. 1-3) and the explanation contributions are computed from the same model and data, but neither quantity is defined in terms of the other, so the association is not true by construction. There is no fitted parameter later renamed as a prediction, and no uniqueness theorem or load-bearing self-citation. The self-citations that appear ([17], [34], [43]) are background or methodological references and do not carry the central argument. The AOPC evaluation uses explanation rankings to define perturbations, but this is a standard sanity check against a random baseline, not a reduction of the claim to its inputs. The most serious weakness is experimental rather than circular: every evaluated model violates distributive fairness, so the claim that procedural-unfairness signals appear 'when distributive fairness is violated' lacks a negative control and is underdetermined. That is a validity threat, not circularity. Under the stated rules, no circular step can be quoted, so the appropriate score is 0.
Assumptions & free parameters
assumptions (6)
- domain assumption The 100 (or 50) instances sampled per group/outcome cell are representative of the full group.
- domain assumption Disparity in aggregated feature attributions across groups indicates procedural unfairness.
- domain assumption AOPC on the Adult dataset is a valid proxy for explanation quality of all methods across other datasets.
- domain assumption Features correlated with the protected attribute are proxies for indirect discrimination.
- domain assumption Standard preprocessing from cited prior work [28,50] is correctly applied.
- standard math Standard definitions of LIME, SHAP, DiCE, and the fairness metrics are assumed as background.
Cite this review
Pith. "Pith review of Explanations as Bias Detectors: A Critical Study of Local Post-hoc XAI Methods for Fairness Exploration." pith.science (2026). https://pith.science/paper/VPTWSSNA
@misc{pith2026250500802,
author = {Pith},
title = {Pith review of: Explanations as Bias Detectors: A Critical Study of Local Post-hoc XAI Methods for Fairness Exploration},
year = {2026},
howpublished = {\url{https://pith.science/paper/VPTWSSNA}},
note = {Machine review of arXiv:2505.00802}
}
read the original abstract
As Artificial Intelligence (AI) is increasingly used in areas that significantly impact human lives, concerns about fairness and transparency have grown, especially regarding their impact on protected groups. Recently, the intersection of explainability and fairness has emerged as an important area to promote responsible AI systems. This paper explores how explainability methods can be leveraged to detect and interpret unfairness. We propose a pipeline that integrates local post-hoc explanation methods to derive fairness-related insights. During the pipeline design, we identify and address critical questions arising from the use of explanations as bias detectors such as the relationship between distributive and procedural fairness, the effect of removing the protected attribute, the consistency and quality of results across different explanation methods, the impact of various aggregation strategies of local explanations on group fairness evaluations, and the overall trustworthiness of explanations as bias detectors. Our results show the potential of explanation methods used for fairness while highlighting the need to carefully consider the aforementioned critical aspects.
Figures
Figures from the paper (19 more)
Reference graph
Works this paper leans on
-
[1]
Amina Adadi and Mohammed Berrada. 2018. Peeking inside the black-box: a survey on explainable artificial intelligence (XAI). IEEE access 6 (2018), 52138–52160
2018
-
[2]
Guilherme Alves, Vaishnavi Bhargava, Miguel Couceiro, and Amedeo Napoli. 2020. Making ML Models Fairer Through Explanations: The Case of LimeOut. In Analysis of Images, Social Networks and Texts - 9th International Conference, AIST 2020, Skolkovo, Moscow, Russia, October 15-16, 2020, Revised Selected Papers (Lecture Notes in Computer Science, Vol. 12602) ...
work page 2020
-
[3]
Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, Adrien Bennetot, Siham Tabik, Alberto Barbado, Salvador García, Sergio Gil-López, Daniel Molina, Richard Benjamins, et al . 2020. Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Information fusion 58 (2020), 82–115
2020
-
[4]
Tom Begley, Tobias Schwedes, Christopher Frye, and Ilya Feige. 2020. Explainability for fair machine learning. arXiv:2010.07389 [cs.LG]
arXiv 2020
-
[5]
Eshta Bhardwaj, Harshit Gujral, Siyi Wu, Ciara Zogheib, Tegan Maharaj, and Christoph Becker. 2024. Machine learning data practices through a data curation lens: An evaluation framework. In The 2024 ACM Conference on Fairness, Accountability, and Transparency . 1055–1067
work page 2024
-
[6]
Vaishnavi Bhargava, Miguel Couceiro, and Amedeo Napoli. 2020. LimeOut: An Ensemble Approach to Improve Process Fairness. In ECML PKDD 2020 Workshops - Workshops of the European Conference on Machine Learning and Knowledge Discovery in Databases (ECML PKDD 2020): SoGood 2020, PDFL 2020, MLCS 2020, NFMCP 2020, DINA 2020, EDML 2020, XKDD 2020 and INRA 2020, ...
work page 2020
-
[7]
Francesco Bodria, Fosca Giannotti, Riccardo Guidotti, Francesca Naretto, Dino Pedreschi, and Salvatore Rinzivillo. 2023. Benchmarking and survey of explanation methods for black box models. Data Mining and Knowledge Discovery 37, 5 (2023), 1719–1778
work page 2023
-
[8]
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. Advances in neural information processing systems 29 (2016)
2016
Show all 49 references
-
[9]
Stephen Casper, Carson Ezell, Charlotte Siegmann, Noam Kolt, Taylor Lynn Curtis, Benjamin Bucknall, Andreas Haupt, Kevin Wei, Jérémy Scheurer, Marius Hobbhahn, et al. 2024. Black-box access is insufficient for rigorous ai audits. In The 2024 ACM Conference on Fairness, Account...
2024
-
[10]
Simon Caton and Christian Haas. 2024. Fairness in machine learning: A survey. Comput. Surveys 56, 7 (2024), 1–38
2024
-
[11]
Juliana Cesaro and Fábio Gagliardi Cozman. 2019. Measuring Unfairness Through Game-Theoretic Interpretability. In Machine Learning and Knowledge Discovery in Databases - International Workshops of ECML PKDD 2019, Würzburg, Germany, September 16-20, 2019, Proceedings, Part I (C...
2019
-
[12]
H Cramer. 1946. Mathematical methods of statistics, Princeton, 1946. Math Rev (Math-SciNet) MR16588 Zentralblatt MATH 63 (1946), 300
1946
-
[13]
Luca Deck, Jakob Schoeffer, Maria De-Arteaga, and Niklas Kühl. 2024. A Critical Survey on Fairness Benefits of Explainable AI. In The 2024 ACM Conference on Fairness, Accountability, and Transparency . 1579–1595
2024
-
[14]
Frances Ding, Moritz Hardt, John Miller, and Ludwig Schmidt. 2021. Retiring adult: New datasets for fair machine learning. Advances in neural information processing systems 34 (2021), 6478–6490
2021
-
[15]
Rudresh Dwivedi, Devam Dave, Het Naik, Smiti Singhal, Rana Omer, Pankesh Patel, Bin Qian, Zhenyu Wen, Tejal Shah, Graham Morgan, et al. 2023. Explainable AI (XAI): Core ideas, techniques, and solutions. Comput. Surveys 55, 9 (2023), 1–33
2023
-
[16]
Alessandro Fabris, Nina Baranowska, Matthew J Dennis, David Graus, Philipp Hacker, Jorge Saldivar, Frederik Zuiderveen Borgesius, and Asia J Biega. [n. d.]. Fairness and bias in algorithmic hiring: A multidisciplinary survey. ACM Transactions on Intelligent Systems and Technol...
-
[17]
Christos Fragkathoulas, Vasiliki Papanikou, Danae Pla Karidi, and Evaggelia Pitoura. 2024. On Explaining Unfairness: An Overview. In 2024 IEEE 40th International Conference on Data Engineering Workshops (ICDEW) . IEEE, 226–236
2024
-
[18]
Sofie Goethals, David Martens, and Toon Calders. 2023. PreCoF: counterfactual explanations for fairness. Machine Learning (2023), 1–32
2023
-
[19]
Przemyslaw A Grabowicz, Nicholas Perello, and Aarshee Mishra. 2022. Marrying fairness and explainability in supervised learning. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency . 1905–1916
2022
-
[20]
Gummadi, and Adrian Weller
Nina Grgic-Hlaca, Muhammad Bilal Zafar, Krishna P. Gummadi, and Adrian Weller. 2016. The Case for Process Fairness in Learning: Feature Selection for Fair Decision Making. In NIPS Symposium on Machine Learning and the Law
2016
-
[21]
Gummadi, and Adrian Weller
Nina Grgic-Hlaca, Muhammad Bilal Zafar, Krishna P. Gummadi, and Adrian Weller. 2018. Beyond Distributive Fairness in Algorithmic Decision Making: Feature Selection for Procedurally Fair Learning. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (...
2018
-
[22]
Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and Dino Pedreschi. 2018. A survey of methods for explaining black box models. ACM computing surveys (CSUR) 51, 5 (2018), 1–42
2018
-
[23]
Victor Guyomard, Françoise Fessant, Thomas Guyet, Tassadit Bouadi, and Alexandre Termier. 2023. Generating robust counterfactual explanations. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, 394–409
2023
-
[24]
Isha Hameed, Samuel Sharpe, Daniel Barcklow, Justin Au-Yeung, Sahil Verma, Jocelyn Huang, Brian Barr, and C Bayan Bruss. 2022. Based-xai: Breaking ablation studies down for explainable artificial intelligence. arXiv preprint arXiv:2207.05566 (2022)
2022 arXiv
-
[25]
Peter Hase, Harry Xie, and Mohit Bansal. 2021. The out-of-distribution problem in explainability and search methods for feature importance explanations. Advances in neural information processing systems 34 (2021), 3650–3666
2021
-
[27]
Aditya Jain, Manish Ravula, and Joydeep Ghosh. 2020. Biased Models Have Biased Explanations. CoRR abs/2012.10986 (2020). https://arxiv.org/abs/2012.10986
2020 arXiv
-
[28]
Faisal Kamiran and Toon Calders. 2012. Data preprocessing techniques for classification without discrimination. Knowledge and information systems 33, 1 (2012), 1–33
2012
-
[29]
Kentaro Kanamori, Takuya Takagi, Ken Kobayashi, and Yuichi Ike. 2022. Counterfactual explanation trees: Transparent and consistent actionable recourse with decision trees. In International Conference on Artificial Intelligence and Statistics . PMLR, 1846–1870
2022
-
[30]
Loukas Kavouras, Konstantinos Tsopelas, Giorgos Giannopoulos, Dimitris Sacharidis, Eleni Psaroudaki, Nikolaos Theologitis, Dimitrios Rontogiannis, Dimitris Fotakis, and Ioannis Emiris. 2023. Fairness Aware Counterfactuals for Subgroups. arXiv preprint arXiv:2306.14978 (2023)
2023 arXiv
-
[31]
Brendan Kennedy, Xisen Jin, Aida Mostafazadeh Davani, Morteza Dehghani, and Xiang Ren. 2020. Contextualizing hate speech classifiers with post-hoc explanation. arXiv preprint arXiv:2005.02439 (2020)
2020 arXiv
-
[32]
Niklas Koenen and Marvin N Wright. 2024. Toward understanding the disagreement problem in neural network feature attribution. In World Conference on Explainable Artificial Intelligence. Springer, 247–269
2024
-
[33]
Jeff Larson, Surya Mattu, Lauren Kirchner, and Julia Angwin. 2016. How we analyzed the COMPAS recidivism algorithm. ProPublica (5
2016
-
[34]
Tai Le Quy, Arjun Roy, Vasileios Iosifidis, Wenbin Zhang, and Eirini Ntoutsi. 2022. A survey on datasets for fairness-aware machine learning. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery 12, 3 (2022), e1452
2022
-
[35]
Dan Ley, Saumitra Mishra, and Daniele Magazzeni. 2023. GLOBE-CE: A Translation-Based Approach for Global Counterfactual Explanations. arXiv preprint arXiv:2305.17021 (2023)
2023 arXiv
-
[36]
Scott M Lundberg and Su-In Lee. 2017. A unified approach to interpreting model predictions. Advances in neural information processing systems 30 (2017)
2017
-
[37]
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. A survey on bias and fairness in machine learning. ACM computing surveys (CSUR) 54, 6 (2021), 1–35
2021
-
[38]
Brent Mittelstadt, Chris Russell, and Sandra Wachter. 2019. Explaining explanations in AI. In Proceedings of the conference on fairness, accountability, and transparency. 279–288
2019
-
[39]
Christoph Molnar. 2020. Interpretable machine learning. Lulu. com
2020
-
[40]
Ramaravind K Mothilal, Amit Sharma, and Chenhao Tan. 2020. Explaining machine learning classifiers through diverse counterfactual explanations. In Proceedings of the 2020 conference on fairness, accountability, and transparency . 607–617
2020
-
[41]
Emmanouil Panagiotou, Manuel Heurich, Tim Landgraf, and Eirini Ntoutsi. 2024. TABCF: Counterfactual Explanations for Tabular Data Using a Transformer-Based VAE. In Proceedings of the 5th ACM International Conference on AI in Finance . 274–282
2024
-
[42]
Pedregosa, G
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011. Scikit-learn: Machine Learning in Python. Journal of Machine Le...
2011
-
[43]
Evaggelia Pitoura, Kostas Stefanidis, and Georgia Koutrika. 2022. Fairness in rankings and recommendations: an overview. The VLDB Journal (2022), 1–28. Explanations as Bias Detectors: A Critical Study of Local Post-hoc XAI Methods for Fairness Exploration • 17
2022
-
[44]
Atul Rawal, James McCoy, Danda B Rawat, Brian M Sadler, and Robert St Amant. 2021. Recent advances in trustworthy explainable artificial intelligence: Status, challenges, and perspectives. IEEE Transactions on Artificial Intelligence 3, 6 (2021), 852–866
2021
-
[45]
Kaivalya Rawal and Himabindu Lakkaraju. 2020. Beyond individualized recourse: Interpretable and interactive summaries of actionable recourses. Advances in Neural Information Processing Systems 33 (2020), 12187–12198
2020
-
[46]
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining . 1135–1144
2016
-
[47]
Dylan Slack, Anna Hilgard, Himabindu Lakkaraju, and Sameer Singh. 2021. Counterfactual explanations can be manipulated. Advances in neural information processing systems 34 (2021), 62–75
2021
-
[48]
Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju. 2020. Fooling lime and shap: Adversarial attacks on post hoc explanation methods. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society . 180–186
2020
-
[49]
Ilse van der Linden, Hinda Haned, and Evangelos Kanoulas. 2019. Global Aggregations of Local Explanations for Black Box models. CoRR abs/1907.03039 (2019). http://arxiv.org/abs/1907.03039
2019 arXiv
-
[50]
Brian Hu Zhang, Blake Lemoine, and Margaret Mitchell. 2018. Mitigating unwanted biases with adversarial learning. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society . 335–340. A Appendix A.1 Datasets Table 5. Dataset description with features, feature de...
2018
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.